Thesis
Voice agents are shifting from command-based interfaces to real-time conversational systems. Early voice assistants were designed for discrete tasks: setting timers, playing music, or checking the weather. Full-duplex voice models, such as NVIDIA's PersonaPlex released in January 2026, listen and respond simultaneously, handle interruptions, and adapt to emotional cues mid-conversation.
Adoption trends have reflected this shift. Voice agent usage on Speechmatics' platform grew 9x in 2025; 15% of the 400 North American business leaders Deepgram surveyed in 2025 said their organizations were actively developing voice AI agents, and one estimate projects the voice AI agents market will reach $47.5 billion by 2034. However, as voice agents move from commands to full-duplex conversation, the bottleneck has moved from model architecture to high-quality training data.
The amount of available data is an important limitation. A 2024 Meta AI paper found that combining all significant spoken dialogue datasets would yield only about 3K hours of dialogue data. A full-duplex model has to listen continuously because backchannels and overlapping speech can occur at any point in a conversation, so it needs training audio that captures both speakers at once. Recording each speaker on a separate microphone channel keeps the two voices apart where they overlap, which a single mixed channel cannot do.
David AI describes itself as an audio data research company. It designs, records, and licenses channel-separated conversational audio datasets for training speech-to-speech (STS), multilingual, and voice interaction models. Existing players tend to either sell from broad pre-collected audio catalogs or build and label datasets to a customer's specification. David AI records participants on independent microphone channels, capturing the turn-taking, interruptions, and overlapping speech that full-duplex systems need.
David AI identifies the conversational behaviors voice models still struggle to replicate and builds the dataset designed to teach them how to do so. Its featured datasets cover two-speaker English conversation, a multilingual set spanning 15+ languages, multi-speaker conversations, and expert conversations across professional domains. David AI chooses which conversational behavior to record and, the company says, tests each dataset against model training outcomes before scaling collection to thousands of hours.
Founding Story
David AI was founded in July 2024 by Tomer Cohen (CEO) and Ben Wiley (CTO), former Scale AI employees, after concluding that voice AI would be the next horizon in AI models.
Cohen studied computer science and economics at Brown University. He spent three years at McKinsey as a senior business analyst before joining Scale AI in 2022 as a product operations lead on generative AI data projects for foundation model labs. Cohen became chief of staff in 2023 and co-founded Scale AI’s Enterprise AI business unit.
Wiley earned a bachelor's degree in computer science from the University of Minnesota in 2021 and joined Microsoft as a software engineer. He then joined Scale AI as a founding engineer and technical lead on Donovan, Scale AI's generative AI platform for government agencies, and became an engineering manager in January 2024.
Wiley and Cohen bonded over their shared interest in multimodal AI, specifically voice AI. OpenAI previewed Voice Engine in March 2024, a model that could recreate a person's voice from a 15-second recording. Meanwhile, voice AI funding was growing, having grown 6.7x from $315 million in 2022 to $2.1 billion in 2024. Wiley and Cohen believed the next wave of voice AI models would need to learn turn-taking and overlapping speech, and that the multi-channel speech datasets most cited in research were, as they later wrote, "dated and only hundreds of hours." The limiting factor would become training-ready conversational audio captured and labeled in the right "shape."
Cohen and Wiley applied to Y Combinator while still at Scale AI, with nothing more than an idea, and were accepted into the Summer 2024 batch. During the batch, they built a phone-call app over a weekend and used it to produce a first small dataset. David AI's first customer came during YC, when it converted that early collection into a $1K contract with a robotics company. By the end of the batch, the company had closed its first six-figure contract with a large AI lab.
Product
Product Philosophy
David AI sells access to proprietary, research-grade audio datasets designed for training STS, multilingual, and voice interaction models. David AI describes a six-step pipeline governing how it identifies, builds, and ships datasets, running from hypothesize, design, and experiment to evaluate and iterate, productionize, and release.
The loop begins with a hypothesis about which audio AI capability is underdeveloped due to insufficient training data. The research team designs the "shape" of data required to teach that capability, specifying key parameters that determine the focus and output of the data. A targeted collection experiment builds the initial sample. The company describes its approach as "research-driven," evaluating each dataset for its "efficacy in model training" as well as its quality, and scaling it "with extreme attention to detail." Iterations continue until the dataset meets internal thresholds for signal quality and model impact. Only then does David AI scale collection to thousands of hours and release the dataset to customers.
To capture audio data, David AI either pays contributors to record natural conversations or purchases datasets. It recruits contributors through Babel Audio, a platform David AI operates, which listed 40K contributors as of May 2026. Each participant records on a separate microphone channel, producing speaker-separated tracks. This is a deliberate design choice to ensure clean separation of voice at the hardware level during conversations. David AI collects audio either through a video-telephony application similar to Zoom or in person.
Featured Datasets

Source: David AI
David AI's homepage features four datasets, each targeting a specific training need in voice AI development, named Converse, Atlas, Chorus, and Dialog. The company also offers additional proprietary datasets not listed on the site, and designs new datasets with customers.
Converse is the flagship English dataset: channel-separated, natural two-speaker conversations across a range of topics. Conversations are composed of unscripted, open-ended dialogue where speakers record on independent microphone channels, producing full-duplex, speaker-separated data. One unverified estimate puts Converse at over 15K hours, against the roughly 3K hours of dialogue a 2024 Meta AI paper found across all significant spoken dialogue datasets combined.
Atlas extends the same channel-separated format across 15+ languages, with detailed metadata on dialects and accents. One customer reports that the languages include Portuguese, Russian, and other European and Indian languages. This dataset shape matters because audio data cannot be translated as straightforwardly as text. Accents, prosody, and dialect carry information that machine translation does not preserve. For example, a model trained on Castilian Spanish can underperform on Mexican Spanish because acoustic and conversational patterns are specific to each variant.
Chorus captures conversations with three or more speakers and was originally designed for training speaker-separation and diarization models. Multi-speaker audio is harder to work with than two-speaker dialogue, because overlapping speech, rapid turn-taking, and uneven speaker volumes degrade single-channel recordings. Diarization, identifying who spoke when, matters for applications such as meeting transcription and conference call analysis.
Dialog is a collection of expert conversations across a range of professional domains. Generic speech models often struggle with domain-specific terminology, conversational norms, and the structured question-and-answer patterns common in professional settings. Dialog addresses this with conversations between domain experts.
Research and Evaluations
In 2026, David AI launched a separate research site that publishes evaluations of voice models. In July 2026, it published a report testing whether large audio language models can stand in for human raters when judging STS models. In September 2026, it released DAI-S2S-ST, David AI's STS leaderboard, which ranks models on human preference and is built on 153K comparative ratings from 283 raters across 819 prompts and seven models. Raters answer 12 rubric questions that cluster into humanness and technical quality, and the prompts are recorded by David AI's own US contributor network and kept private to prevent models from training on them. Because the ratings come from David AI's own contributors, the leaderboard also demonstrates the kind of human-preference data the company can collect for customers.
Market
Customer
David AI does not publicly disclose individual customer names. In its Series A announcement in May 2025, David AI said it had grown to "support most of the 'Mag 7' companies and nearly every leading audio AI lab." By its Series B in October 2025, it described itself, in narrower terms, as "working with several of the Mag7 and most of the leading AI labs." A May 2025 profile characterized the company as having "found a voracious customer base among some of the biggest names in tech." Its disclosed customers are therefore a small set of very large buyers. David AI's 2026 job postings state that it has "brought on most FAANG companies and AI labs as customers."
Overall, David AI’s target customers fall into three segments. The first is AI research labs building STS and voice interaction systems, which the company says are "bottlenecked by access to training data and evaluations." The second is large technology companies building voice interfaces into consumer products; David AI's homepage says its datasets are used by "Fortune 100 companies and research labs." The third is companies building voice into physical products, with the Series B announcement naming humanoid robots, wearables, and personal assistants among the uses of its data.
Market Size
David AI operates within two expanding markets, voice AI and AI training data. The global speech and voice recognition market was valued at $19.1 billion in 2025 and is projected to reach $104.1 billion by 2034, a 20.3% CAGR from 2026 to 2034. The AI training dataset market was worth $3.2 billion in 2025 and is expected to reach $16.3 billion by 2033, a 22.6% CAGR from 2026 to 2033. The downstream markets that consume this data are expanding faster. One estimate projects the voice AI agents market will grow from $2.4 billion in 2024 to $47.5 billion by 2034, a 34.8% CAGR from 2025 to 2034.
David AI's addressable opportunity sits within the audio and speech segment of the AI training dataset market. Image and video data held the largest share of that market in 2025, at 41.9%, while audio is expected to record the highest growth rate of any segment through 2032. One report sizes the AI training dataset market at $6.1 billion in 2025 and puts audio at 25% of that market, implying $1.5 billion of audio training data spending. The specific niche David AI serves, conversational channel-separated speech for training STS models, is a smaller and earlier-stage subset of that total.
Competition
Scale AI: Founded in 2016 by Alexandr Wang and Lucy Guo, Scale AI is a data engine designed to generate bespoke datasets at operational scale. As of September 2026, Scale AI had raised $15.9 billion in total funding, including a $1 billion Series F in May 2024 led by Accel at a $13.8 billion valuation, with participation from NVIDIA and Y Combinator. In June 2025, Meta invested $14.3 billion for 49% of Scale AI as a non-voting stake, in a deal that valued the company at more than $29 billion. As part of the deal, Wang joined Meta as its first Chief AI Officer to lead Meta Superintelligence Labs, and Scale AI's chief strategy officer Jason Droege became interim CEO. Droege had dropped the interim title by January 2026, signing the company's outlook for the year as CEO. In July 2025, Scale AI cut 14% of its staff, largely in its data-labeling business.
In Scale AI's customer-driven business model, the customer defines the dataset it needs through the data engine, and Scale AI builds and labels it to specification with its contractor network. Scale AI also produces reinforcement learning from human feedback (RLHF) and evaluation data, including preference, safety, and model-evaluation workflows, to improve model performance. Its competitive advantage is built on delivering datasets, evaluation data, and post-training integration to customer specifications. Although Scale AI's network of over 240K contractors in 2025 was about six times the 40K contributors on David AI's Babel Audio platform in May 2026, channel-separated conversational audio recorded at the hardware level is harder to procure. David AI specializes in this niche, having built a collection operation optimized for it.
Defined.ai: Defined.ai, originally named DefinedCrowd, was founded in 2015. Founder and CEO Daniela Braga holds a PhD in speech technology from the University of A Coruña and a master's degree in applied linguistics from the University of Minho. As of September 2026, the company had raised a total of $78.6 million in funding, including an $11.8 million Series A in July 2018 led by Evolution Equity Partners and a $50.5 million Series B in May 2020. The company reported 65% year-over-year revenue growth in 2025.
Defined.ai describes itself as "the world's largest AI marketplace," offering pre-collected and structured training datasets across text, voice, and image modalities. Its speech catalog spans spontaneous dialogue, scripted monologue, and spontaneous interactive voice response categories. Customers access this data through Defined.ai's marketplace and Neevo, a crowdsourcing platform with over 1.5 million contributors as of September 2026.
Defined.ai is the closest dataset competitor to David AI, since its spontaneous dialogue datasets, multilingual locale coverage, and domain-specific call-center-style datasets compete with David AI's Converse, Atlas, and Dialog datasets. Some customers, however, describe Defined.ai's marketplace as more difficult to use than David AI, citing its API integrations as more error-prone. Another customer noted that Defined.ai's datasets focus on "scale and vision of overall speech" and less on quality than David AI's.
Protege: Founded in 2024, Protege is a data platform that licenses real-world data from data providers, including media content, audio recordings, and de-identified health records, and prepares it for AI training. In January 2026, Protege raised a $30 million Series A extension led by Andreessen Horowitz. It has raised a total of $65 million as of September 2026. Its audio and speech catalog includes conversational audio in Indic languages and Arabic dialects, call center conversations, and recordings of vocal bursts and interruptions, and the company works with "the majority of the 'Magnificent Seven,'" the same buyers David AI sells to. Protege's catalog is built from data licensed from third parties, while David AI pays contributors to record its core datasets.
Magic Data: Magic Data takes a similar catalog approach to Defined.ai but pairs it with annotation tooling and managed services. QingQing Zhang founded the company in 2016 after earning a PhD at the Chinese Academy of Sciences' Institute of Acoustics, where she later worked as an associate research fellow, and completing postdoctoral work at the French National Centre for Scientific Research. Magic Data has grown to over 200 employees as of September 2026. It raised a pre-Series A in 2017, a Series A in 2018, a 2019 extension, and a Series B in May 2021 led by Vantron Capital.
Magic Data positions itself as "a global leading multi-modal conversational AI data services provider" and reported over 200K hours of training data, including conversational and speech data spanning Asian languages, English dialects, and European languages, in January 2022. Beyond its catalog, Magic Data launched Annotator 5.0, an AI-assisted data annotation and management platform, in 2021, and sells custom dataset delivery and collection services. The annotation software differentiates it from David AI, which sells datasets rather than tooling.
The competitive dynamic between Magic Data and David AI comes down to breadth versus depth. Magic Data is optimized for procurement convenience, with large catalogs, contributor networks, and bundled services that get buyers usable data quickly. David AI takes a narrower approach, with four featured datasets, each designed for training readiness around specific conversational behaviors.
Panels: Panels is a newer entrant competing directly with David AI in conversational audio data for voice AI training. Aaron Wenk and Jason Le, who both studied operations research and financial engineering at Princeton, founded Panels in 2025 and joined YC's Summer 2025 batch, which came with YC's standard $500K investment. At its August 2025 launch, Panels reported over 10K vetted voice contributors covering over 20 languages and over 100 countries.
Panels' offerings include speaker-separated conversational audio, single-speaker scripted recordings, turn-taking evaluation, and custom datasets built to a customer's design. David AI's advantage comes from knowing which dataset to build. Panels is competing on how quickly it can build any dataset a customer needs.
The common thread across these competitors is scale, in the form of more hours, more contributors, and more services. David AI's differentiation is its dataset feedback loop. The company evaluates datasets for model training efficacy before scaling collection, so what it learns about which dataset structures improve which model capabilities informs what it collects next.
Business Model
David AI generates revenue by designing, collecting, and curating conversational audio data, then licensing it to AI labs and large technology companies. The model is problem-centric, meaning David AI chooses what to build rather than waiting for a customer's specification. Its research team decides which datasets are valuable to create based on market-wide training requirements, builds them to a quality standard, and publishes them for licensing. Beyond its featured catalog, David AI also offers custom dataset design, in which customers work with David AI to decide what the data should capture.
The core commercial motion is B2B data licensing. Customers request samples, sign a data license agreement, and receive off-the-shelf datasets within one to two days. Pricing is tied to specific use cases, dataset scope, licensing length, and the number of users who need access. Collection is the main cost, covering recruiting speakers, running recording sessions, and scaling hours. David AI pays its contributors to record original speech, and roles on its Babel Audio platform paid $24 to $28 per hour as of May 2026. Because each dataset is recorded rather than scraped, the cost of a dataset grows with the hours recorded.
Traction
David AI's revenue grew quickly after its founding. By January 2025, six months after it started, David AI had exceeded seven figures in revenue, working with customers ranging from the largest technology companies to startups. At its Series A in May 2025, David AI had crossed eight figures in annual revenue run rate. The ramp from a $1K prototype contract during YC's Summer 2024 batch to an eight-figure run rate took less than a year. David AI has stated that its data "has already been used to train several of the best speech models on the market," though it has declined to name clients. By May 2025, David AI's corpus held over 100K hours of audio across 15+ languages. The company had about 50 employees as of July 2026, up from about 15 a few months earlier.
Valuation
As of September 2026, David AI had raised $80.5 million in total funding. Its last round was a $50 million Series B in October 2025, led by Meritech Capital with participation from NVIDIA and existing investors, at a reported $500 million valuation. That valuation was about four times the value set at its $25 million Series A.
Alt Capital and Amplify Partners co-led the Series A round, which valued David AI at more than $100 million in May 2025. Because eight figures means at least $10 million, the $500 million valuation was at most 50x the eight-figure run rate David AI disclosed in May 2025, the most recent revenue figure it has published. Before that, First Round Capital led a $5 million seed round in January 2025, with participation from Y Combinator, and Y Combinator invested $500K at pre-seed in September 2024.
Key Opportunities
Expanding the Dataset Catalog Through Targeted R&D
David AI's research team can target dataset formats the market needs next, such as emotional prosody, noisy environments, far-field microphones, or additional languages. Each new dataset widens David AI's catalog, which already includes proprietary datasets beyond the four it features. As David AI accumulates signal on which dataset structures improve model performance, a new entrant would have to rebuild that knowledge from scratch. Atlas, the company's multilingual dataset, is a particularly large expansion surface. With 15+ languages already covered and dialect and accent metadata built into collection, David AI can extend that coverage as voice AI adoption spreads beyond English-speaking markets.
Selling Evaluations Alongside Training Data
David AI has begun to move into evaluation. Its Series B announcement framed the company's work around both training data and evaluations, and its DAI-S2S-ST leaderboard ranks STS models from OpenAI, Google, xAI, Amazon, and Alibaba. Panels already lists turn-taking evaluation among its offerings, so the category is contested from the start.
Scale AI is a reference point, having started as a data-labeling company and expanded into evaluation, RLHF, and benchmarking, including public leaderboards of frontier models. David AI could extend its leaderboard into repeatable benchmarks for latency tolerance, turn-taking accuracy, interruption handling, diarization quality, and multilingual robustness, sold alongside its training data. Selling both the dataset and the framework that measures a model trained on it would embed David AI deeper in a customer's development pipeline.
Supplying Data for Full-Duplex Models
The industry is shifting from traditional pipelines that chain speech recognition, a language model, and text-to-speech toward full-duplex models that listen and speak simultaneously. NVIDIA's PersonaPlex team described the gap directly: "Traditional systems … let you customize the voice and role, but conversations feel robotic with awkward pauses, no interruptions, and unnatural turn-taking."
Full-duplex models need to learn when to pause, interrupt, and backchannel as well as what to say, and those behaviors require structured, channel-separated conversational data. NVIDIA trained PersonaPlex on 1.2K hours of real conversations from the Fisher English corpus, supplemented with synthetic dialogues, because real conversations "present varied natural interaction patterns that current TTS systems cannot simulate reliably."
Key Risks
AI Labs Can Vertically Integrate
David AI's pitch rests on two assumptions: (1) high-quality conversational audio is scarce, and (2) it is operationally difficult to produce. Both could erode as AI labs build internal expertise.
ElevenLabs, an AI audio and speech company, said employees at over 60% of Fortune 500 companies had adopted its tools as of its January 2025 Series C. When ElevenLabs raised a $500 million Series D at an $11 billion valuation in February 2026, it said it would expand its ElevenAgents voice-agent platform and its research on emotional conversational models. A company that runs enterprise voice agents and trains its own conversational models could build the conversational corpus David AI sells rather than license it.
Google, Meta, and Apple each run consumer voice products, and Meta has already brought a data vendor partly in-house, paying $14.3 billion for 49% of Scale AI in June 2025. If a Mag 7 customer decides to build a private internal audio corpus, David AI's licensing relationship with it could become redundant.
Speech AI companies add build-versus-buy pressure. One unverified report notes that Deepgram processes over 50K years of audio internally and that Speechmatics uses self-supervised learning to reduce its reliance on human-labeled data. Customers who already process large volumes of audio may decide it is cheaper to build their own data pipeline than to keep licensing from an external vendor like David AI.
Proving ROI at Renewal
Customers expect measurable model improvements from datasets. Isolating the impact of any single dataset is hard to do in practice. A voice model's performance reflects pre-training data, fine-tuning data, architecture choices, hyperparameter tuning, RLHF, and post-training evaluation all at once. If a customer trains on David AI's Converse alongside three other datasets while changing its model architecture and adjusting hyperparameters in the same cycle, there is no clean way to attribute a specific share of improvement to any one dataset. Procurement teams may push back on premium pricing they cannot tie to measured performance gains.
David AI's contract structure adds to this pressure. Revenue comes from licensing agreements that vary in length, and one customer reports that the price rises at renewal as a dataset grows. If a customer that signed a longer-term deal cannot show that the data improved its model by the time renewal comes around, it has leverage to renegotiate or move to a cheaper competitor. Because David AI's disclosed customers are a small set of very large buyers, a single lost renewal would weigh heavily on revenue.
Data Vendor Regulatory Compliance
David AI's exposure is the consent behind its recordings. Illinois' Biometric Information Privacy Act treats voiceprints as biometric identifiers and requires a written release from the subject before a private company collects or purchases one, and David AI's Babel Audio platform listed 40K contributors as of May 2026. David AI also purchases some datasets, whose provenance it has to verify rather than control. The US Copyright Office's May 2025 pre-publication report concluded that some uses of copyrighted works for AI training will qualify as fair use and some will not, leaving the status of purchased third-party audio unsettled. If buyers require vendors to document consent and provenance for every recording, compliance overhead could slow enterprise deals.
Summary
Voice AI is entering a phase where natural conversation, rather than transcription accuracy, sets the bar. As systems move toward real-time, full-duplex interaction, the supply of conversational training data falls short, and a 2024 paper counted about 3K hours across all significant spoken dialogue datasets combined. Most vendors optimize for volume or breadth rather than the interaction dynamics these models require.
David AI treats this as an R&D problem. Founded in 2024 by two former Scale AI employees, the company designs, collects, and licenses studio-grade conversational datasets structured for STS, multilingual, and voice interaction training. David AI crossed eight figures in annual revenue run rate by May 2025, had raised $80.5 million as of September 2026, and was valued at a reported $500 million in its October 2025 Series B, led by Meritech Capital with participation from NVIDIA. The key question is whether David AI can keep defining the next dataset format voice models need, and prove that its data measurably improves them, before its largest customers decide to build that data themselves.



