"Open source models just raise the floor of what people are interested in... open source models just mean that nobody's buying anything that Kim K3 can already do." - Oswald Nitski [01:04]
"We end every week with so much more money in the bank like the business is very healthy and we can't spend money fast enough to service all of the demand that we have today." - Oswald Nitski [00:04]
Disclaimer: Orignal content owned by or sourced from third parties. It does not represent the views of 'Nuggets' platform or it's team. AI is used extensively across this platform including for summaries. Accuracy is not guaranteed, there can be mistakes. Any info or content on this platform is not a financial, legal, or investment advice. Do your own research. Refer for complete disclosures:- Terms of Use · Full Disclaimer
"There's a whole category of latent demand that people aren't even... trying to do with models yet... long horizon tasks like setting up a procurement agent to fully automate your procurement team for months on end." - Oswald Nitski [02:52]
"I don't think there's an ROI problem right now. I think we're in a period of exploration and experimentation where there's more tolerance, more patience to get that ROI calculation right now." - Oswald Nitski [07:51]
"I'm very careful never to delegate judgment or decision-making to models because... they make you think that it's doing the right thing, but you have to be paranoid with them still." - Oswald Nitski [26:44]
"Spend on Mercor directly translates to more revenue for our customers." - Oswald Nitski [43:53]
Speakers & Credentials
Harry Stebbings (Host): Founder of 20VC (The Twenty Minute VC), a leading venture capital podcast interviewing prominent founders, tech executives, and investors.
Oswald Nitski (Guest): Chief Product Officer (CPO) at Mercor, an AI-powered talent marketplace and evaluation platform connecting domain experts with frontier AI labs to produce high-quality evaluation and training data.
1. Executive Summary
Open Source Model Floor Elevation: Open-source models (like Moonshot AI's Kimi or open-weights alternatives) do not cannibalize Mercor's business; rather, they raise the baseline capability floor, forcing frontier labs to acquire richer evaluation and training datasets to unlock non-commodity capabilities [01:40].
Unlocking Latent Demand: Contrary to claims that 90% of enterprise workflows are solvable today, immense latent demand remains untapped in long-horizon, multi-step tasks (e.g., autonomous, months-long procurement agents) where current top models score around 50% on benchmark suites like APEX [02:52], [05:27].
Enterprise Sensitivity & Deployment: Organizations comfortably put non-core functions (HR, procurement) on closed proprietary APIs, but guard their core IP tightly, often leveraging local or open-weights deployments to maintain total data custody [04:02], [05:00].
Paradigm Shift in Software PM Ratios: Coding agents like Claude and Cursor are shrinking engineering bottlenecks, triggering a structural shift toward higher PM-to-Engineer ratios where product managers must focus heavily on business impact and high-level judgment over tool mastery [13:21], [30:32].
The ROI Horizon: Enterprises do not face an immediate AI ROI crisis; instead, they operate in a grace period of aggressive exploration where spending is viewed as essential to maintain technological parity and capture growth [07:51], [09:15].
Data Frontier Shifts to RL Environments: Demand in data labeling has evolved rapidly from Supervised Fine-Tuning (SFT) and Preference Ranking (RLHF) to complex Reinforcement Learning (RL) simulation environments and multi-file world states [41:19], [42:16].
Financial Position & Hypergrowth: Mercor experiences extreme cash-flow generation, expanding headcount by more than 10x in a year and ending every week with net positive cash reserves despite aggressive reinvestment to meet insatiable lab demand [11:54], [37:36].
Rise of RL in Cyber and Physical Data: Adversarial categories like cybersecurity and physical robotics data represent uncapped, continuous reward functions where static performance standards do not apply, creating durable long-term data demand [51:12], [58:46].
The rapid rise of open-weights models—such as Moonshot AI's Kimi series [01:25]—has sparked debate over whether open-source AI will cannibalize the commercial data provisioning market. Open models raise the baseline floor of commoditized intelligence, making it redundant to sell training or evaluation data for basic, solved tasks [02:02]. However, rather than shrinking the market, open source forces frontier labs to purchase hyper-specialized evaluation and fine-tuning datasets that target capability gaps beyond the current floor [01:40].
Claims that 90% of enterprise workflows can already be handled by standard open models reflect a narrow view of current user prompts rather than total potential demand [02:22]. A vast category of "latent demand" remains untouched because enterprises have not yet attempted to deploy models for complex, long-horizon tasks—such as an autonomous procurement agent that operates independently for months with minimal check-ins [02:52]. On benchmark suites like Mercor’s APEX, top frontier models currently score around 50% on long-horizon, multi-step tasks, proving significant room for growth [05:27].
Enterprise Security, Sufficiency, & Model Specialization
Enterprise willingness to share data with frontier model providers varies based on how core the workflow is to competitive differentiation [03:53]. Commodity tasks like HR administration or procurement are frequently routed through proprietary, closed model APIs [04:11]. Conversely, highly sensitive intellectual property—such as specialized legal arguments or medical advisory memos—faces strict privacy constraints [04:25]. In these scenarios, companies deploy open-weights models inside private clouds or local infrastructure to maintain complete data custody [04:54].
Workflows fall into two distinct operational categories:
Sufficiency-Based Workflows: Binary completion tasks where a model reaches maximum utility once completed (e.g., updating fields in a CRM) [05:43].
Continuous Uncapped Reward Categories: Workflows where quality can continually improve without a fixed ceiling (e.g., drafting complex legal arguments or diagnosing medical cases) [06:03].
This distinction supports the thesis that enterprises will require bespoke, domain-specific models tailored to their internal operational goals—whether prioritizing aggressive revenue growth, margin expansion, or regulatory compliance [06:23].
Token Economics, ROI Grace Periods, & Software PM Evolution
Enterprises are not experiencing an immediate "AI ROI crisis" [07:51]. Instead, organizations operate in an exploratory phase characterized by high tolerance for experimental spend while token prices and model capabilities evolve [08:01]. Budgeting approaches differ markedly across companies: ClickHouse multiplied its infrastructure spend by 6x to maintain a performance edge at the frontier [08:38], whereas organizations like Uber, Microsoft, and X have experimented with per-head allocations for end-user coding agents [08:43].
At fast-growing companies like Mercor, monthly AI token spend already exceeds total developer payroll [11:36]. This shift alters team structures:
Higher PM-to-Engineer Ratios: Because coding agents dramatically reduce engineering bottlenecks, software teams require fewer engineers per product manager [30:32].
Elimination of "Skill Issues": Syntax mastery and manual boilerplate writing are no longer primary constraints; product success depends on high-level business judgment, experimental design, and systems architecture [13:41], [25:03].
Product Surface Area Simplification: The ability to generate thousands of lines of code instantaneously creates the risk of bloat; PMs must actively strip away unnecessary features to maintain product simplicity [12:22], [12:55].
Shift Away from Specialized Tools: Teams increasingly bypass dedicated UI mockup platforms like Figma in favor of rapid cloud-based prototyping and direct code rendering [13:53].
The Evolution of Data Provisioning: RL Environments & Cyber Realities
The human data market has undergone rapid technical transitions since the launch of InstructGPT, moving sequentially from Supervised Fine-Tuning (SFT) to Preference Ranking (RLHF), and now to Reinforcement Learning (RL) simulation environments [15:37], [40:41].
[ Supervised Fine-Tuning (SFT) ] ──> [ Preference Ranking (RLHF) ] ──> [ RL Environments & World States ]
RL environments represent the current frontier for training autonomous agents [41:19]. Rather than processing static text prompts, agents interact with simulated software applications (e.g., deep Salesforce mocks) alongside complex initial "world states"—multi-file system snapshots representing thousands of local computer files [41:46].
A rapidly expanding category in data creation is cybersecurity [50:23]. Cyber offense and defense operate as an adversarial game with uncapped reward functions, where static datasets quickly become obsolete [51:12]. Furthermore, while small boutique data vendors often rely on VC-subsidized, founder-led manual annotation, frontier labs require infrastructure capable of operating at massive scale with strict quality guarantees [44:42].
The Reference Vault
4. Data & Figures
Data Point
Value
Context
Timestamp
APEX Long-Horizon Benchmark Score
~50%
Current score of top frontier models on Mercor's APEX benchmark for multi-step tasks
Latent Demand in AI Capability Expansion [02:52]: Most market analysis incorrectly projects future AI utility based on current user prompts and visible API traffic. In reality, a massive latent demand curve exists for tasks users do not yet attempt because current models lack the necessary reliability. As frontier capabilities expand to handle long-horizon tasks (e.g., multi-month procurement automation), entirely new software categories will emerge.
Sufficiency vs. Uncapped Reward Functions [05:38]: Workflows naturally divide into two structural types. Sufficiency-based workflows reach a binary performance cap once completed accurately (e.g., updating a CRM). Uncapped reward workflows continuously yield higher returns with incremental performance improvements (e.g., legal strategy, medical diagnostics, cybersecurity defense). Data provisioning efforts focus increasingly on these uncapped domains, where performance ceilings do not exist.
The High-Agency PM / Engineer Paradigm [13:41]: As automated coding tools compress the time required to turn software designs into working code, manual implementation is no longer the primary bottleneck in product engineering. Consequently, product managers must evolve into high-agency business operators who emphasize experimental rigor, statistical validity, and systemic simplification over feature volume.
The Knowledge Dissemination Lag in AI Services [19:04]: Enterprise demand for forward-deployed engineering (FDE) and specialized AI integration services stems from a regional concentration of AI talent in San Francisco. In the short term, enterprises must hire external services to deploy complex agentic architectures. Over time, as this knowledge distributes globally, AI agent setup will integrate directly into standard internal software engineering roles.
6. Anecdotes
The Annotation Platform Surface Area Chaos [15:24]: Early in Mercor's development, the team attempted to support every customized human data format requested by clients—spanning supervised fine-tuning, preference ranking, and multimodal data. Managing hundreds of distinct ad-hoc annotation setups simultaneously created extreme operational complexity. This forced leadership to institute strict operational guardrails and focus exclusively on high-demand, repeatable data categories.
The High-Agency Surf & Sauna Retreat [53:56]: Mercor organized an offsite for its annotation platform team in Tofino, Canada—a remote coastal town known for cold-water surfing. The team spent the trip taking surf lessons and using floating saunas. Oswald highlighted this trip to illustrate how shared physical activities help maintain high agency, intense focus, and strong interpersonal alignment within fast-moving startup environments.
The 15-Minute Robot Water Retrieval Demo [01:00:18]: Harry Stebbings recounted watching a live demo where a domestic robot took 15 minutes to retrieve a bottle of water from a kitchen fridge. Unbeknownst to the audience, the robot was being teleoperated by Mercor CEO Brandon Wang from an adjacent room. Stebbings used this story to question current robotics capabilities, while Oswald drew parallels to the early days of autonomous vehicles, predicting a similar eventual breakthrough.
7. References & Recommendations
Companies & Platforms
Mercor: AI talent marketplace and data platform connecting domain experts with frontier model labs [00:14].
ClickHouse: Open-source database management company that scaled its AI token budget by 6x to maintain performance at the frontier [08:38].
Salesforce: Global enterprise software provider reported to spend $300M annually on Anthropic LLM tokens [10:12].
Fireworks AI: Specialized AI inference and model hosting platform referenced during discussions on enterprise model diversification [06:17].
Figma: Collaborative interface design platform being increasingly bypassed in favor of cloud-based design and live code prototyping [00:00], [14:01].
Palantir: Enterprise software firm known for forward-deployed engineering services [18:48].
Surge AI: Human data provider for frontier AI models, referenced as a competitor [43:03], [57:02].
Waymo: Autonomous driving subsidiary of Alphabet, cited as proof of long-term success in scaling physical AI deployments [01:01:05].
AI Models & Benchmarks
Moonshot AI (Kimi / Kim K3): Open-weights frontier model series developed in China [01:25], [02:09].
InstructGPT: Early benchmark model from OpenAI that established human feedback (RLHF) techniques [15:37], [42:37].
APEX Benchmark: Mercor's internal evaluation benchmark designed to test models on long-horizon, complex operational workflows [05:27].
People
Brandon Wang: Co-founder and CEO of Mercor [01:09], [11:23].
Alex Karp: CEO of Palantir, cited for his commentary on enterprise skepticism toward sharing data with model providers [03:32].
Lynn Xiao: Founder of Fireworks AI, referenced for her thesis on enterprise-specific model specialization [06:17].
Marc Benioff: CEO of Salesforce, referenced regarding enterprise token expenditures [10:12].
Matan Grinberg: Founder of Factory, cited for his view that forward-deployed services can reflect underlying product limitations [19:50].
Geographic Locations
Tofino, Canada: Coastal town on Vancouver Island, destination for Mercor's team retreat [53:56].
Jul 25, 2026
Anthropic Engineer on How to Get the Most Out of Claude Code | Thariq Shihipar | 16 Jul 2026 | South Park Commons
1. Executive Briefing TL;DR Capability Overhang & Exponential Growth: Frontier AI models are currently more intelligent than existing UI, interaction harnesses, and human prompting mental models permit them to display; freezing model devel…
>10x
Headcount growth at Mercor over the past 12 months