your system language is:English

Scaling AI Inference: Base 10’s Path to 30x Growth

Cover

📺 Today’s recommended deep-dive video: https://www.youtube.com/watch?v=XAbKflCncDo


The $1 Billion Inference Race: Scaling Custom AI at the Edge of Capacity

Tuhin Srivastava, CEO of Base 10, explains how his company achieved 30x growth in a single year by providing the backbone for the next generation of AI-native applications. He breaks down why custom models are winning over vanilla open-source and how the global compute shortage is forcing a radical reimagining of infrastructure.

Core Question: How can companies maintain hyper-growth in a market defined by extreme compute scarcity and a shift toward highly specialized, vertical AI models?

Highlights

  • Custom models account for 95% of inference traffic compared to off-the-shelf weights.
  • The “supply crunch” is worse than reported, with clusters running at mid-90s utilization.
  • Vertical AI moats are built on proprietary “user signals” and deep workflow integrations.
  • Operational “pager culture” is the hidden requirement for surviving the infrastructure wars.

⏱️ Reading time: approx. 6 minutes · Saves you about 37 minutes vs. watching.

Want to take notes while watching? Click the image below and let AI Notebook capture the key points for you 👇

AI Notebook


The Sovereignty of Custom Models

Why Vertical AI Wins the Moat War

The existential threat from frontier labs is countered by the unique user signals and deep workflow integrations that only specialized application layers can capture.

Consider a medical scribe like Abridge; it succeeds by embedding itself into specific clinician behaviors that a general-purpose model from OpenAI cannot see. By owning the data generated from physician edits, these companies can post-train models on a unique reward signal, effectively building a moat through specialized intelligence that serves a vertical market better than any generic frontier tool ever could.

This move toward specialization explains why nearly all traffic on Base 10 involves custom weights rather than off-the-shelf open source. Every major scale customer is optimizing for specific performance, latency, and quality metrics that require fine-tuning.

💡 Digging Deeper

Q: Why do application companies exist if frontier models are so good?
A: Labs lack access to the specific “user signal” found in deep workflows, like how a doctor interacts with an EMR.
Q: What percentage of your customers use vanilla models?
A: Less than 5%; almost everyone makes modifications for quality or performance before hitting scale.
Q: Is enterprise adoption already peaked?
A: No, 99% of the market hasn’t come online yet; we are still serving the “AI native” early adopters.


Navigating the Compute Crunch

The Reality of 90% Utilization

While the media often discusses a supply crunch, the reality on the ground is far more severe than most outside observers realize.

Base 10 maintains mid-90s utilization across its global footprint, leaving almost no slack for sudden spikes in demand. To navigate this, the company has built a runtime fabric that spans 18 different clouds and 90 clusters worldwide. This abstraction layer allows them to onboard a new provider in a foreign country in less than half a day, ensuring they can capture supply wherever it exists while maintaining enterprise-grade reliability and failover capabilities for their customers.

Scaling at this speed requires avoiding “grifty” suppliers who lack the operational experience to manage high-stakes data centers.

A flowchart showing the Base 10 fabric: User Request -> Global Load Balancer -> 18 Managed Clouds/90 Clusters -> Runtime Fabric -> Inference Output.


Geopolitics and the Multi-Chip Future

The DeepSeek Moment and NVIDIA’s Moat

The rise of high-quality Chinese models like DeepSeek represents a massive economic subsidy for American enterprises willing to utilize their low-cost intelligence.

Srivastava argues that missing out on these low-cost, high-frontier models would be a strategic loss for the US innovation ecosystem. If a model can be network-bound and secured, its origin matters less than the value it provides. Furthermore, the world is moving toward a multi-chip reality where inference-specific hardware will handle decodes while NVIDIA continues to dominate the training and high-speed ecosystem.

NVIDIA remains the gold standard because of the massive developer ecosystem surrounding CUDA, making it the only choice for companies that need to move fast.

A concept map showing the NVIDIA Moat: Center node is "NVIDIA Dominance," branching out to "CUDA Ecosystem," "Supply Chain Excellence," "Developer Mindshare," and "Rapid Iteration Speed."


Key Takeaways

The shift from generalized AI to specialized agents is the defining trend of the current market. As companies prove product-market fit with frontier models, they immediately move toward post-training and custom inference to drive down costs and improve accuracy. This “loop” between inference and post-training is where the most significant value will be created over the next five years.

Reliability in this environment is not just a software problem but an operational one. The infrastructure companies that survive will be those that embrace a “pager culture,” treating even minor system latencies as P0 emergencies. As intelligence becomes a utility, the ability to manage working capital and secure multi-year compute contracts will separate the billion-dollar platforms from the commodity providers.


Q&A

Q1: How much did Base 10 grow in the last year?
A: The company grew 30x in terms of scale and is approaching a billion-dollar revenue run rate.

Q2: What is the “Jevons Paradox” in the context of AI?
A: It is the idea that as inference becomes cheaper and more efficient, total demand actually increases because developers find more ways to embed intelligence into workflows.

Q3: Are Chinese models like DeepSeek a security risk?
A: If models are network-bound and data is handled correctly, they are seen more as a high-quality, low-cost “subsidy” for US innovation.

Q4: Why is Base 10 in 18 different clouds?
A: To abstract away the supply crunch and provide a single runtime fabric that can failover across different providers and geographies.

Q5: What is the most important cultural trait for an infrastructure company?
A: An “operations culture” where leadership and engineering are deeply connected to the stability of the system, often via a literal pager.

Q6: How long are the contract lengths for the newest chips?
A: For high-end hardware like the B200, suppliers are increasingly demanding 3-to-5-year contracts with significant upfront prepayments.

Q7: What comes after AGI?
A: In a world where intelligence is solved, the only thing left is inference—the act of applying that intelligence to specific problems.

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts