Fish Audio Secures $50M to Expand AI Voice Generation for Creators and Enterprises
Fish Audio, a voice synthesis startup with 8M users and $21M ARR, closes seed round led by Coreline Ventures to scale customizable voice models across creative and business applications.
Last verified:
The Funding Round
Fish Audio, a voice synthesis startup founded by former NVIDIA researcher Shijia Liao, announced a $50M seed investment on July 28, according to TechCrunch. The round was anchored by Coreline Ventures and Capital Today, with additional backing from 359 Capital, Parable, Play Time, Alphalist Partners, Bayhouse Ventures, Carya Venture Partners, and HF0. The funding values the Palo Alto-based company at a reported $200M–$300M post-money valuation, though exact terms remain undisclosed.
The capital injection arrives as the startup has already reached measurable traction: 8 million monthly active users across its open-source and commercial platforms, paired with $21M in annual recurring revenue. This revenue multiple underscores investor confidence in Fish Audio’s market positioning at the intersection of creator tools and enterprise infrastructure.
From Open-Source to Enterprise Scale
Fish Audio’s trajectory began with a single developer effort. Liao, dissatisfied with the expressiveness of available synthetic voices, trained a voice generation model on consumer GPU hardware and released it to GitHub under an open-source license. The Fish Speech repository has accumulated over 31,000 stars, attracting indie developers, game studios, and short-form video creators seeking alternatives to flat, robotic voice synthesis.
Over the past 12 months, Fish Audio has released five distinct models—four for speech generation and one for speech-to-text. Three of the speech generation models remain open-source; the newest iteration, the S2.1 Pro variant, is available exclusively through Fish Audio’s paid API tier. This tiered approach allows the startup to capture value from both the hobbyist ecosystem (via open releases) and commercial deployments.
The company’s control surface—more than 15,000 natural language parameters—differentiates its offering. According to Fish Audio CEO and co-founder Rissa Cao, this granularity addresses divergent use cases. Avatar platforms like HeyGen prioritize naturalness and lip-sync fidelity; gaming studios demand expressive character voices; voice-agent providers like LiveKit need low-latency, conversation-grade synthesis. A single generic model cannot satisfy all three.
Voice Consent and Takedown Automation
In mid-2026, Fish Audio faced public criticism after allegations surfaced that user-submitted voice samples had been added to its training library without explicit consent. The startup’s original DMCA takedown process, while available, proved slow—a friction point that risked creator trust and regulatory scrutiny around voice rights.
Cao disclosed to TechCrunch that Fish Audio has now automated its takedown workflow. Creators can upload a voice clip or contractual proof of ownership; Fish Audio commits to removing matching voices from its platform within 3 minutes. While this accelerates dispute resolution, the underlying challenge remains: preventing non-consenting uploads in the first place. Automation addresses the symptom (slow removal) rather than the root cause (initial upload without verification).
Why This Matters
The $50M raise validates a market thesis: synthetic voice quality and control are now table-stakes for AI applications spanning entertainment, customer service, and gaming. Fish Audio’s $21M ARR—on an estimated $250M post-money valuation—implies a 1.2x revenue multiple, typical for fast-growing B2B SaaS, but generous for a 12-month-old commercial product. This suggests investors are pricing in near-term TAM expansion as enterprises migrate from legacy voice APIs to controllable synthesis.
The consent automation signals an emerging operational standard. Startups built on user-generated content—including voice, video, and text—will face increasing pressure to implement rapid dispute resolution. Fish Audio’s 3-minute SLA may become table-stakes rather than a differentiator. Regulatory attention to voice rights (particularly in the EU under emerging AI Act precedent) will likely force further investment in provenance tracking and verification, not just reactive removal.
Frequently Asked Questions
What makes Fish Audio's voice models different from existing text-to-speech tools?
Fish Audio offers over 15,000 natural language controls enabling fine-grained expressiveness and steering, compared to generic TTS. Its models are optimized for both creative use cases (character voices, video content) and enterprise applications (customer support, sales automation).
How did Fish Audio respond to the voice consent controversy?
CEO Rissa Cao told TechCrunch the company automated its DMCA takedown process. Creators can now submit voice samples or contracts to prove ownership; matching voices are removed from the platform within 3 minutes.
Who are Fish Audio's paying customers?
According to TechCrunch, HeyGen (AI avatar platform), Sanas (speech enhancement), and Plaud (voice AI device) are already using Fish Audio's enterprise APIs.