Thesis
In November 2022, OpenAI released ChatGPT, which became the fastest-growing consumer product of all time. Three years later, it had 800 million weekly active users, and OpenAI's annualized revenue had crossed $20 billion. The combined AI capital expenditures from large tech companies have nearly tripled from $162 billion in 2022 to $448 billion in 2025.
Every one of these AI systems depends on data. Language models are trained on terabytes of text; computer vision models on millions of labeled images; AI agents on reward signals from simulated environments. This data must be collected, organized, annotated, and, most importantly, quality-assured. The consequences of low data quality range from ineffective AI (chatbots that hallucinate, coding agents that produce faulty code) to downright harmful AI (medical bots that misdiagnose tumors, self-driving cars that crash). This has given rise to an entire market dedicated to AI data labelling, which was valued at $3.8 billion in 2024 and projected to reach $17.1 billion by 2030, representing a 28.4% CAGR over that period.
Since ChatGPT's release in 2022, breakthroughs in AI research have forced the data labeling industry to adapt. The rise in unsupervised learning, particularly self-supervised learning, a method in which AI is trained on unlabeled, raw data, has reduced AI labs' reliance on large, labeled datasets. Rich Sutton's Bitter Lesson that human expertise is less valuable than raw computing power is generally accepted amongst AI researchers, and suggests that using human experts to label terabytes of training data is ineffective. As a result, the industry has shifted from static pre-training datasets to post-training services like supervised fine-tuning (SFT), reinforcement learning from human feedback (RLHF), and reinforcement learning (RL) environments that train AI agents. At each stage, the role of human data work has changed, but never disappeared.
Surge AI takes a strict quality-first approach, pairing more than 1 million labelers and domain experts with human-in-the-loop workflows. It holds contracts with a majority of the most valuable frontier labs, and has helped train models such as Google's Gemini and Anthropic's Claude. Surge is completely bootstrapped, having raised zero outside capital, a rarity in a market where its competitors have raised billions. Surge has changed its product line at each of those shifts, offering the products that became necessities after every breakthrough. The company is built on the idea that regardless of how AI research advances, humans will still need to be involved in making AI safe, capable, and aligned. Its goal is to "raise AGI with the richness of humanity".
Founding Story
Surge AI was founded by Edwin Chen (CEO) in 2020. Chen grew up in Crystal River, Florida, a town with a population of 3.4K. He took calculus in eighth grade, got a full-ride to the boarding school Choate in Connecticut, and studied math, computer science, and linguistics at MIT. He left MIT in his third year to work at Clarium Capital (Peter Thiel's former hedge fund) and never returned to school. For the next decade, he worked on content moderation and recommendation algorithms at Twitter, Google, and Facebook. At each job, he kept running into the same problem: it was hard to get high-quality, human-labeled data at a large scale.
Inspired by this, he left Twitter in 2020 to start Surge, kickstarting it with "a couple million" of his personal savings. Chen's first engineering hire was Andrew Mauboussin (CTO), a Harvard computer science graduate. Mauboussin had previously led Twitter's Spam and Integrity efforts to safeguard democratic elections and applied machine learning to financial modeling at Kensho Technologies. He joined Surge in August 2020 as its first engineer and has overseen the company's engineering teams that built Surge's data labeling and quality control systems.
Chen and Mauboussin's shared background in content moderation and data quality at Twitter gave them firsthand experience with the difficulty of accurately labeling data at scale. This became the core of Surge's value proposition. The company was profitable from nearly day one, having landed clients like OpenAI in 2021 and Anthropic in 2022. Since then, Surge has held contracts with almost every frontier AI lab. Chen has written that staying independent "lets us prioritize research and rigor over theater and hype."
Product
Generally, there are three big steps to training an AI model: preparing data, training the model (pre-training), and fine-tuning the model (post-training). These three steps have roughly correlated with three major shifts in AI research over the last few years, and following each shift, Surge has provided the types of products that quickly became necessities.
Pre-Training Data
The first shift, from 2020 to 2022, concerned pre-training data. GPT 3.5, the model used for the commercial ChatGPT release, was trained on a 45 TB corpus of data, including books, Wikipedia articles, and internet forums. Five months later, its successor GPT-4 was trained on 1 petabyte of data (22 times as much), including images, higher-quality text, legal documents, and code. GPT-4 was viewed as a substantial improvement over its predecessor, and a trend became clear: more data led to better LLMs.
During this time, Surge provided many of the labeled datasets used to train frontier language models. It employed a human workforce, including domain experts, to label large datasets using their expertise. It also used a human-in-the-loop variation of this approach, where AI generates its own data and labels it, but humans critique its performance.
This has allowed Surge to differentiate its product in many key ways. Surge employs domain experts like "doctors, lawyers, investment bankers, Fields Medalists, Harvard professors, and more", selecting labelers who were at the top of their respective fields. Its data is international (operating in over 70 languages) and multimodal (ranging from images, audio, and video). And it provides off-the-shelf data, allowing labs to purchase labeled datasets, which distill thousands of hours of expert reasoning, for immediate use.
Post-Training Data
The second shift, from 2022 to 2024, was about post-training methods. As OpenAI's o1 chain-of-thought models began outperforming traditional models on key benchmarks, there emerged a greater focus on post-training techniques like supervised fine-tuning (SFT) and reinforcement learning from human feedback (RLHF). These techniques involve using human preferences to tweak an AI's behavior and require a large number of skilled human evaluators.
Surge provided a variety of post-training services. SFT describes training a generalized LLM, which is capable of many things, to do one thing (eg, coding, answering legal questions) really well. This requires specialized, high-quality datasets in specific fields as well as domain experts to assess the AI's responses, which Surge provided. For RLHF, Surge's annotators compared different responses to the same question, recorded which one was better, and gave explanations for why. Its domain experts wrote rubrics and verifiers, which are essentially scorecards for AI performance. And finally, it offered red-teaming and evaluation services, where its annotators tried to break an agent's safety and security guidelines, then gave feedback on how to fix them.

Source: Data Annotation
RL Environments
The third shift, which began in 2024, was marked by a move towards AI agents. As the economic value of AI shifted from chatbots, which require prompts like ChatGPT, to fully agentic automations like Cursor, Claude Cowork, or OpenClaw, so has the data required to train them.
As a result, Surge began offering Reinforcement Learning (RL) environments as a service, and by September 2025 had created an internal organization dedicated to building them. Instead of static datasets, RL environments are essentially reward-laden digital playgrounds. AI agents are trained by roaming around these playgrounds and learning what constitutes good behavior. Surge published research in January 2026 outlining the diagnostic framework that labs use to interpret what RL environments measure. From this, it identified what it calls the "hierarchy of agentic capabilities": tool use, planning and goal formation, adaptability, groundedness, and common-sense reasoning. Even the best-performing models in that study failed roughly 40% of tasks. As of August 2026, RL environments are marketed as Surge's biggest product.
The money behind that shift is concentrated. Anthropic's leaders have discussed spending more than $1 billion on RL environments over a single year, a figure comparable to Surge's entire 2024 revenue.
Benchmarks and Research
Surge also publishes research and creates evaluations intended to show which models perform best on work that resembles real professional tasks. As of August 2026 it maintained eight public benchmarks: Chartography (chart understanding), HANDBOOK.md (long-context instruction following), EnterpriseBench: CoreCraft (agentic workflows), ComplexConstraints (professional instruction following), GDP.pdf (multimodal document reasoning), Riemann-bench (research-level mathematics), Hemingway-bench (writing), and Antidote (real-world professional judgment). In August 2026 it combined all eight into the Tuesday Work Index, a composite measure of professional intelligence.

Source: Surge AI
Industry Data Quality
While the concept of labeling data might seem straightforward, there are many perverse incentives inherent to the practice. Chen has said that industry benchmarks, such as Arena (an online leaderboard where people vote on the best AI responses), optimize for the wrong things. Instead of reading and fact-checking the response, it disproportionately rewards answers that are formatted better or use more emojis. More worryingly, faulty training can cause AI to behave in explicitly manipulative ways. Research published in February 2026 shows that incorrectly executed reinforcement learning from human feedback (RLHF) can cause chatbots to exhibit sycophancy and dishonesty, telling people what they want to hear.
Because of this, data labelers need to have deep knowledge and expertise in their domains. In 2022, Surge relabeled Google's GoEmotions dataset, which classified Reddit posts by emotion. Google had sent posts to workers in India for labeling, but Surge employees familiar with American internet culture found that 30% of the labels were severely wrong. Posts expressing excitement like "LETS F**KING GOOOOO" had been classified as anger, and posts expressing sarcasm like "Yay, cold McDonald's. My favorite" were classified as love. Even simply being familiar with the language isn't enough. For example, an annotator unfamiliar with US politics might rate a comment of "let's go Brandon!" as positive, even though the phrase was widely used as a pejorative dog whistle against former President Joe Biden. These small mistakes, when compounded at a large scale, can have massive impacts on the quality of the AI being trained.
Workforce
To execute its data labeling, Surge employs more than 1 million specialized and non-specialized gig workers from 50 countries using a subsidiary called Data Annotation. Most data labeling competitors use workers from less developed countries in the Global South who are paid $5 an hour. By contrast, as of August 2026 Data Annotation advertised $25 to $50 an hour for generalist work and $75 to $150 an hour for engineering, legal, medical, and finance specialists.

Source: Data Annotation
To screen these workers, Data Annotation makes prospective annotators pass tests mimicking the labeling they would do on the job. Surge also monitors annotators' performance with hidden quizzes, manual review, and machine learning algorithms that optimize for performance. Chen has cited Surge's quality controls and deep technical expertise as the secret to its success.
Market
Customer
Surge sells primarily to frontier AI labs, having held contracts with many since the earliest days of the AI boom. Its clients have included OpenAI (since 2021), Anthropic (since 2022), Google, Meta, Microsoft, and Mistral AI. The buyer set is not exclusively commercial: as of August 2026 its customers also included the US Army and the US Air Force.
This makes its customer base extremely concentrated, which is both a strength and a risk. On the upside, having deep relationships with a small number of labs allows Surge to tailor its workforce, tooling, and quality controls to the specific demands of its customers. Since these labs are high-spending, a single contract can move the company's revenue meaningfully. On the downside, the reverse is also true, because losing even one client impacts a disproportionate share of revenue. Most of these customers also hold contracts with Surge's rivals, so switching costs are low, and labs can shift volume toward whichever vendor delivers the best data.
This puts Surge in constant competition with other data-labeling firms, which all offer broadly similar services to the same narrow set of buyers. Because the pool of frontier labs is small and largely already covered, Surge competes mainly on data quality and turnaround speed rather than on winning untouched accounts.
Notably, Surge does not try to compete on price. It charges 50% to ten times more than competitors. Because of this, it tries to pitch directly to the data scientists at tech firms, in the hopes that they will recognize the quality of Surge's data and be more willing to pay for it. Its Google contract began with a spontaneous two-hour call with one of the search company's researchers and grew to more than $100 million per year.
Market Size
The data labeling market was estimated to be worth $2.6 billion as of 2026, growing at approximately 22% annually to around $7 billion by 2031. Broader reports, which also include data collection in their definitions, have valued the market at $3.8 billion in 2024, projecting it to reach $17.1 billion by 2030. Notably, Surge's $1.2 billion revenue in 2024 alone would represent a substantial share of these estimates, suggesting that these traditional market estimations underestimate the full scope of spending from frontier AI labs on data services.
An alternative way to calculate TAM is to work bottom-up from AI spend. In 2026, the largest US hyperscalers (Microsoft, Alphabet, Amazon, Meta, and Oracle) have CapEx approaching $700 billion. Compute spend accounts for over 50% of AI companies' total expenses, and companies spend an estimated 10-20% of their compute spend on data. Multiplying $700 billion of CapEx by 50% and then by 10-20% produces a data-services TAM of $35 billion to $70 billion. That last input comes from the CEO of Turing, a company that sells data services, so it is directionally useful rather than neutral.
The gap between the top-down market reports ($7 billion to $17 billion) and the bottom-up estimation ($35 billion to $70 billion) likely reflects the fact that frontier labs spend on data through large, bespoke contracts, like Meta's $150 million spend on Surge in 2024, rather than through standardized annotation platforms with public data. The market is also broadening in scope. As the industry shifts from static labeled datasets toward RL environments and agentic evaluations, the definition of "data services" is widening, and spending per project is increasing accordingly.
Whichever framing is used, Surge is large relative to the market it sells into. Its $1.2 billion of 2024 revenue represents between 1.7% and 17% of these TAM estimates, depending on which definition is used.
Competition
Competitive Landscape
The frontier data market has concentrated around a handful of vendors that sell expert human judgment rather than volume annotation. Meta's June 2025 investment in Scale AI pushed several of the largest labs toward Scale's rivals, and a March 2026 breach at Mercor briefly did the same in reverse. Both events moved contracts quickly, which is the defining feature of the category: the buyers are few, the contracts are large, and the switching costs are low.
Two groups compete for that spend. The first is the expert-marketplace group, which recruits credentialed people and routes them to labs: Surge, Mercor, Handshake, Turing, and micro1. The second is a newer set of RL-environment specialists, including Prime Intellect and Mechanize, which never ran a labeling business and are building simulated environments directly. Prime Intellect raised a $130 million Series A in July 2026 and reported over $100 million in annualized revenue. Surge's advantage over the second group is its installed base at the labs; its exposure is that RL environments are a software product, and none of its incumbency in annotation transfers automatically.
Competitors
Scale AI: Founded in 2016, Scale AI provides data labeling and evaluation services for training AI models. It built its own data annotation platforms, Remotasks for computer vision and Outlier for language models. It provided human-in-the-loop annotation, synthetic data generation, multimodal training data, and model evaluation. Scale generated $870 million in revenue in 2024, serving major AI labs such as OpenAI, Google, Meta, and Microsoft, as well as Fortune 500 companies such as General Motors and Time. As of August 2026 it had raised a total of $1.6 billion, including a $1 billion Series F in May 2024 at a $13.8 billion valuation led by Accel. Other notable investors include Dragoneer Investment Group, Greenoaks, and Tiger Global Management.
In June 2025, Meta bought a 49% non-voting stake in Scale AI for $14.3 billion and acquired its CEO, Alexandr Wang, to lead its newly formed Meta Superintelligence Labs. The June 2025 deal valued Scale at $29 billion but immediately raised concerns around Scale's neutrality. OpenAI, Google, Microsoft, and xAI ended their contracts, forming a new market opportunity that was quickly capitalized on by competitors like Mercor and Surge.
Mercor: Mercor was founded in 2023 as an AI-based hiring platform matching workers with opportunities tailored to their skills. After Meta purchased its stake in Scale AI in 2025, Mercor pivoted to matching companies with specialized domain experts to help with AI model training. As of August 2026 it had raised a total of $483 million, including a $350 million Series C in October 2025 at a $10 billion valuation, led by Felicis. Other notable investors include General Catalyst, Menlo Ventures, Alumni Ventures, and SV Angel. It has since broadened its scope to building RL environments. As of August 2026, Mercor is reportedly the fastest-growing company in the category: it hit $2 billion in gross annualized revenue in mid-2026, and in July 2026 was in talks to raise $500 million at a $20 billion valuation.
Handshake: Handshake was founded in 2014 as a college recruiting platform connecting students with employers, and its network reaches 18 million students and alumni across 1.5K university partners. As of August 2026 it had raised a total of $434 million, including a $200 million Series F in January 2022 at a $3.5 billion valuation led by Coatue. Other notable investors include Lightspeed Venture Partners, Kleiner Perkins, Spark Capital, and Reach Capital. Handshake entered the AI data labeling market in 2025, focusing on matching educated, English-speaking workers with post-training data tasks, and created its own annotation platform, Handshake AI. That business reached $1 billion in gross annualized revenue in 2026, making it the clearest example of a recruiting network converting its supply of credentialed people into AI training revenue.
Turing: Turing was founded in 2018 to connect companies with software engineers through its interviewing and AI matching process. In 2022, when OpenAI contracted them for help training its code generation model Codex, Turing pivoted from developer placement to AI training services. Turing reported approximately $300 million in ARR and is reportedly profitable. As of August 2026 it had raised $270 million in total funding, including an $111 million Series E in March 2025 at a $2.2 billion valuation, led by Malaysia's sovereign wealth fund Khazanah Nasional Berhad. Other notable investors include Gaingels, Menlo Ventures, Alumni Ventures, Plug and Play, and Lightspeed Venture Partners. As of 2026, Turing's platform maintained a pool of over 3 million developers globally for data annotation and evaluation, and it has expanded into RL environments and multimodal agent training beyond coding.
micro1: Founded in 2022, micro1 recruits and vets domain experts, accepting a small share of applicants, and routes them to AI labs for model training work. As of August 2026 it had raised a total of $41.6 million, including a $35 million Series A in September 2025 at a $500 million valuation led by 01 Advisors. Its gross annualized run rate grew from $100 million to $500 million in the eight months to August 2026. It is the smallest of the group and the clearest evidence that the expert-marketplace model is replicable on modest capital, which is the part of Surge's position that is hardest to defend.
Business Model
Although Surge doesn't have public pricing information, it operates on a per-project fee model for its labeled datasets and agent feedback. Its contracts are large: Meta spent $150 million in 2024, and Google's contract grew to more than $100 million per year. Its highest cost is annotator labor from its 1 million gig workers, hired through its subsidiary Data Annotation, which has not disclosed information about its finances.
Since these workers are classified as contractors and not employees, Surge is an asset-light business. In practice, this means that Surge avoids the fixed costs that would otherwise come with a million-person workforce, including payroll taxes, benefits, paid leave, severance obligations, office space, and equipment. Its contractors work remotely using their own devices, which helped it secure many annotators during the COVID-19 pandemic. This also means that labor cost scales directly with project volume, because Surge can simply allocate more or fewer annotators depending on project size. This model of converting a fixed cost to a variable cost is a large reason why Surge has operated profitably without ever raising outside capital, and why its margins can absorb its premium pricing strategy.
The tradeoff falls more on the worker side than it does on Surge. Contractors bear the income volatility, lack employment protections, and have limited recourse when task availability changes. These issues have drawn scrutiny to Surge and the data labeling industry as a whole. This also creates legal exposure: as of May 2025, Surge faced a lawsuit accusing it of misclassifying its workers, which the plaintiffs' counsel described as "wage theft on a massive scale". A successful challenge could force it to reclassify its workers as employees, which would remove most of the asset-light advantage.
Surge is secretive and operates under NDAs with its customers. Nothing on the Data Annotation site mentions Surge by name. It didn't respond to questions about its ownership of Data Annotation in 2023, claiming its secrecy comes from clients' demands for confidentiality. In 2025, Chen said the secrecy was because Surge was "too busy to discuss (its) work externally".
Traction
Surge discloses little, so most of what is known comes from third parties. As of 2024, Surge reported $1.2 billion in revenue and 250 employees, including full-time, part-time, and consultants. One of its biggest competitors, Scale AI, employed four times as many people with only $870 million in revenue in the same year. Chen has also said that Surge has been profitable from nearly day one.
That revenue figure made Surge the highest-revenue company in the data labeling space as of 2025, though the ranking has since tightened: Mercor hit $2 billion in gross annualized revenue in mid-2026 and Handshake reached $1 billion, while Surge has published no figure more recent than 2024.
The clearest external evidence of Surge's standing is that its customers cite its evaluations in their own model releases. OpenAI cited GDP.pdf in the GPT-5.6 release in July 2026, and Anthropic cited both GDP.pdf and Riemann-bench in the system card for Fable 5 and Mythos 5 in the same month. For a vendor whose product is judgment, being named as the yardstick by the labs that buy from it is a harder signal to manufacture than a customer logo.
Valuation
Surge has never raised outside capital, so it has no priced round and no valuation set by an investor. In July 2025, it was in talks to raise up to $1 billion, a mix of primary and secondary capital, to let employees sell shares and to capitalize on demand following the exodus of customers from Scale AI. As of September 2025 Forbes put the company's value at an estimated $24 billion, with the raise being discussed at $30 billion. No round had been announced as closed as of August 2026, so every figure here is a talks-stage mark rather than a completed one.
Chen has called most VC-backed Silicon Valley startups a "get-rich-quick scheme", and has said he bootstrapped Surge with his own savings from Big Tech. Chen owns 75% of Surge, a stake worth an estimated $18 billion, which made him at age 37 the youngest member of the 2025 Forbes 400. Against that $24 billion mark, Surge's 2024 revenue of $1.2 billion implies a revenue multiple of roughly 20x. Mercor's $20 billion discussion against $2 billion of gross annualized revenue implies a comparable multiple on a company that is growing faster and disclosing more.
Key Opportunities
Capturing RL Environment Spend as Budgets Concentrate
Anthropic's leaders have discussed spending more than $1 billion on RL environments over a single year, and Surge stood up an internal organization dedicated to building them by September 2025. Surge builds these environments differently from its competitors: instead of designing them top-down, its domain experts create one coherent world organically, populating it with realistic entities and tasks from their own experience. That allows Surge to expose which traits models fail at, then sell the environment expansion that targets that exact weakness. Its January 2026 research on the hierarchy of agentic capabilities is the diagnostic that makes that upsell legible to a buyer.
Converting Benchmark Adoption Into Purchasing
OpenAI cited Surge's GDP.pdf benchmark in the GPT-5.6 release and Anthropic cited GDP.pdf and Riemann-bench in the Fable 5 and Mythos 5 system card, both in July 2026. When a lab adopts an external benchmark as a reporting standard, the vendor that built it owns the definition of the gap and is the natural supplier of the training data that closes it. Surge combined all eight of its benchmarks into the Tuesday Work Index in August 2026, which extends that position from individual evaluations to a single index labs are invited to be measured against.
Key Risks
Model Improvement
There have been concerns that AI evaluation will become harder as models improve beyond both general and expert human intelligence. Surge has begun experimenting with workflows where an AI assists a human in evaluating another AI, collaborating with Anthropic researchers on a proof-of-concept, but that is a response rather than a resolution. If the marginal value of a human evaluator falls faster than Surge can move up the capability curve, its premium over cheaper vendors compresses at exactly the moment its labor cost advantage matters most. Surge's own research is a partial hedge, since it measures where models still fail, but a benchmark that models keep beating is a shrinking business.
Vertical Integration
More companies have begun moving their data annotation in-house. Uber dropped Scale AI following its deal with Meta and started labeling its own data, and Surge faces that same threat. Cohere, one of Surge's earliest customers that initially saw "big lifts" in performance from its data, dropped Surge and moved its data annotation in-house in 2023. Meta's Llama 4 model relied heavily on AI to create and label its own synthetic data, without the need for Surge's workforce. OpenAI, a customer since 2021, dropped Surge in 2025 but still holds contracts with competitors Mercor and Invisible, which suggests the loss was a vendor choice rather than a shift away from bought data. Losing contracts like these is especially serious for Surge, since it hasn't raised any external funding to cushion the blow.
China Exposure
As of August 2026, the same suppliers selling human-preference data to OpenAI and Anthropic were also supplying Chinese labs, and Surge was among them, with Chen travelling to China to meet lab executives directly. The top six Chinese AI labs spend about $500 million a year with American data labeling companies, and the trade runs without an export gate. That sits awkwardly against Surge's US government work, since its customers have also included the US Army and the US Air Force, and Scale AI's Alexandr Wang has publicly criticized suppliers that serve both. If training data is brought under the kind of export controls already applied to chips, a revenue line becomes a compliance problem and a defense customer becomes a conflict.
Data Security
Two documents meant for Surge's gig workers were leaked in July 2025. One was a Google Doc made public by mistake, which covered safety guidelines for training chatbots on sensitive issues such as medical advice, sexually explicit content, violence, and hate speech. Another was a spreadsheet listing which sites to use and which to avoid when training Anthropic's chatbot, Claude, so it would sound "helpful, honest, and harmless". These leaked documents raised concerns from clients like Anthropic about data security, an issue that became more acute after Mercor suffered a large-scale breach in 2026. In a category where switching costs are low, a security incident is one of the few things that reliably moves a contract.
Legal Issues
As of May 2025, Surge faced a class action lawsuit in San Francisco Superior Court for misclassifying full-time workers as independent contractors, meaning it didn't have to provide vacations and health insurance. The suit, brought by the Clarkson Law Firm, names Chen personally as a defendant and cites "millions" in unpaid wages. It also claims that Surge's actions are part of a broader trend that will continue, as "tech giants race to dominate the AI space". Surge has called the lawsuit "without merit" and continued defending itself in the ongoing case.
Summary
Surge AI sells expert human judgment to frontier AI labs. It was founded by a former Big Tech research scientist who left MIT in his third year and spent a decade running into the same problem: high-quality labeled data was hard to get at scale. Since 2020, it has adapted its product after every AI breakthrough, be that changes in pre-training, post-training, or a shift towards AI agents. As of 2026, its products span data annotation, SFT and RLHF, RL environments, and a public benchmark suite that its own customers cite in their model releases.
Surge reached $1.2 billion in revenue in 2024 with a small team and no outside capital, and it has published nothing more recent. In the same window, Mercor reached $2 billion in gross annualized revenue and Handshake $1 billion, both on venture money Surge declined to take. The question ahead is whether a bootstrapped, secretive, quality-obsessed model holds as customers weigh building in-house, better-capitalized rivals compound faster, and the definition of the data industry keeps shifting toward environments Surge has to build rather than people it can recruit.



