your system language is:English

Scaling AI Inference: BaseTen CEO on 30x Growth and Compute

Cover

📺 Today’s recommended deep-dive video: https://www.youtube.com/watch?v=XAbKflCncDo


The Last Market: Scaling AI Inference in a World of Constrained Compute

AI is transitioning from a experimental novelty to a ubiquitous layer of intelligence embedded in every business workflow. Tuhin Srivastava, CEO of BaseTen, explains how his company grew 30x in a single year by providing the specialized infrastructure that powers this shift.

Core Question: How can companies navigate extreme compute scarcity and model customization to build defensible AI applications at scale?

Highlights

  • Why 95% of BaseTen’s tokens come from custom, post-trained models rather than “vanilla” open source.
  • The “supply crunch” is as much about operational talent and data center reliability as it is about silicon availability.
  • How the Jevons Paradox is driving a shift from simple chatbots to long-running, “concierge” agentic workflows.
  • The strategic necessity of the “pager culture” in maintaining the high-utilization clusters required for modern AI.

⏱️ Reading time: approx. 6 minutes · Saves you about 37 minutes vs. watching.

Want to take notes while watching? Click the image below and let AI Notebook capture the key points for you 👇

AI Notebook


The Shift to Specialized Intelligence

Beyond Off-the-Shelf Models

Companies are no longer satisfied with generic, off-the-shelf models; they are aggressively moving toward specialized intelligence that encodes their unique user signals.

Tuhin notes that while the majority of the market today is driven by AI-native startups like Abridge or Decagon, the massive wave of enterprise adoption is still just starting to crest. These companies are transitioning from simple API usage to in-housing their intelligence through sophisticated RL techniques and post-training, allowing them to create specialized workflows that are impossible for frontier labs to replicate because the labs lack access to specific vertical data.

This shift signifies a maturation of the AI stack where performance, latency, and reliability become the primary drivers of model selection over raw brand recognition or general-purpose benchmarks. BaseTen’s acquisition of a research team specialized in post-training highlights the deep link between how a model is trained and how it performs during inference. By integrating these two traditionally separate stages, developers can optimize their weights for specific hardware, ensuring that their proprietary data translates into a tangible competitive advantage in production environments where every millisecond counts.

A flowchart showing the 'Inference-Post-training Loop': User signal enters the system -> Model generates output -> Evaluation (Evals) identifies gaps -> Post-training refines model weights -> Optimized model is redeployed for Inference.

💡 Digging Deeper

Q: Is the independent application layer at risk from frontier labs like OpenAI?
A: No, because labs lack access to the specific “user signal” found in deep workflows, like a physician’s edits to medical notes.

Q: Why are Chinese open-source models like DeepSeek gaining traction in the US?
A: They offer frontier-level capability at roughly 20% of the cost of closed-source alternatives, making them economically irresistible for high-volume tasks.

Q: What is the primary reason for customizing a model today?
A: While quality is the starting point, the ultimate goal is performance optimization—making models “better, faster, and cheaper” for a specific domain.


The Infrastructure Bottleneck

Managing Global GPU Scarcity

The supply crunch isn’t just about a lack of chips; it is a fundamental shortage of operators who can reliably manage high-performance data centers at scale.

BaseTen currently operates across eighteen different clouds and ninety clusters worldwide, maintaining utilization rates in the mid-nineties to meet surging demand. This fragmented landscape requires a unified runtime fabric that can abstract away the underlying complexity of diverse hardware providers while maintaining the strict service level agreements that mission-critical enterprise applications demand.

Buying capacity is becoming a high-stakes financial game where three-to-five-year contracts with significant upfront payments are the new standard for accessing next-generation hardware like Nvidia’s B200s. In this environment, the cost of capital becomes a strategic lever just as important as software engineering. Companies that can bridge the gap between volatile demand and rigid hardware supply are the ones that will ultimately control the distribution of intelligence across the global economy. Owning the compute is no longer just a utility—it is a strategic asset.

💡 Digging Deeper

Q: How fast can BaseTen spin up a new provider?
A: Due to their unified runtime fabric, they can integrate a new data center provider into their global cluster in about half a day.

Q: Will there be a multi-chip world beyond Nvidia?
A: While diversification is desirable, Nvidia’s supply chain excellence and the Cuda ecosystem make them nearly impossible to unseat in the short term.

Q: How has the term length for compute contracts changed?
A: It has shifted from short-term “pay-as-you-go” to mandatory 3-to-5-year commitments for high-end GPUs.


Culture and the Future of Agency

The “Concierge Everything” Economy

In the world of infrastructure, the pager is the ultimate arbiter of truth; if the system fails, someone must be awake to fix it.

Tuhin describes a “hero culture” rejection where flat leadership and first-principles thinking are the only ways to survive 30x annual growth. This operational rigor is necessary because as the cost of inference drops due to Jevons Paradox, users don’t save money—they simply consume significantly more intelligence through longer-running, more complex agentic workflows.

The future of the application layer looks like “concierge everything,” where every consumer has a personalized team of agents for health, education, and life management. This shift from selling software seats to selling units of cognition marks the beginning of the “last market”—inference. Even if artificial general intelligence is eventually achieved, the physical act of serving that intelligence to a user remains the core bottleneck and the most valuable service a company can provide. We aren’t building less software; we are building more software, faster, with intelligence at the center.

A concept map showing the transition from 'SaaS' (Software as a Service) to 'IaaS' (Intelligence as a Service). SaaS features: Seats, Subscriptions, Tools. IaaS features: Units of Cognition, Agentic Workflows, Personalized Concierge.

💡 Digging Deeper

Q: What is the “Pager Culture” Tuhin mentions?
A: It is an operations-first mindset where everyone, including leadership, is tethered to system reliability, treating infrastructure downtime as a P0 crisis.

Q: How does the Jevons Paradox apply to AI?
A: Making intelligence cheaper doesn’t reduce spending; it causes developers to insert “a hell of a lot more intelligence” into their apps, driving total consumption up.

Q: Does AI mean fewer software engineers?
A: Tuhin argues we will just build a “ton more software” and more complex tools rather than reducing the headcount of builders.


Key Takeaways

Inference has emerged as the “terminal market” for the AI era. While training creates the model, inference is where the actual economic value is delivered to the end-user. As models become more specialized through post-training, the infrastructure required to run them must become more sophisticated, moving away from commodity “GPUs-as-a-service” toward intelligent runtime fabrics that can manage complex, multi-cloud deployments.

The competitive moat for AI companies is moving from the model itself to the proprietary user signals and workflows that the model powers. By owning the inference layer and the post-training loop, companies can create a flywheel where every user interaction makes the model more specialized and efficient. This creates a “concierge” experience for consumers that is both highly personalized and economically defensible for the provider.

Finally, the physical reality of compute cannot be ignored. In a supply-constrained world, operational excellence in data center management and strategic access to hardware are the primary constraints on growth. The winners in this space will be those who can marry high-level software abstraction with the gritty, “pager-wearing” reality of global hardware operations.


Q&A

Q1: What percentage of BaseTen’s workload is currently based on custom models?
A1: Over 95% of the tokens served are for custom models where the client has modified the weights or optimized the runtime for their specific use case.

Q2: Is the GPU supply crunch getting better?
A2: No; while more chips are being produced, the market is “supplier and operationally crunched.” There is very little “slack” compute, with clusters often running at 90%+ utilization.

Q3: How does BaseTen handle reliability across 18 different clouds?
A3: They built a unified runtime fabric that spans all clusters, allowing for seamless failover and abstraction of the underlying hardware complexity from the developer.

Q4: Why is post-training becoming so important for inference companies?
A4: Post-training and inference are two sides of the same coin. How you train a model (such as quantization-aware training) directly dictates how efficiently it can be served.

Q5: What is the biggest challenge in hiring for a hyper-growth AI company?
A5: Finding “first-principles” people who thrive in a flat, high-accountability environment and who understand the “pager culture” of infrastructure.

Q6: What happens to software pricing as we move to agentic workflows?
A6: Pricing is shifting from “seats” (digitalization) to “units of cognition” (intelligence), where companies sell the successful completion of a task rather than just the tool.

Q7: Will specialized inference chips replace Nvidia GPUs?
A7: There will be decode-specific and inference-specific chips, but Nvidia’s ecosystem and supply chain speed make them the dominant choice for any company needing to move fast today.


Leave a Reply

Your email address will not be published. Required fields are marked *

Related Posts