跳到正文
Modal 官方工程博客·· 2 天前AI 评分43

TypeSafe AI 在 Modal 上三天内将 Jev 扩展至超万亿 token

TypeSafe AI scales Jev to trillions of tokens in three days on Modal

AI 导读

TypeSafe AI 的首个模型 Jev 在 Modal 上于发布后三天内处理超 1 万亿 token,并在 OpenRouter 1K-10K token 上下文榜单登顶,承接 15.1% 的请求。Jev 并非大语言模型,而是非自回归的 System One Model,以状态和问题为输入,并行输出带概率的结构化类型值,训练采用自研的 RLCD 方法。

正文

TypeSafe AI’s manifesto is contrarian for an AI lab: “Build Prod, Not God.” Their first model, Jev, delivers fast, structured decisions with frontier intelligence. In the course of a launch week, demand for Jev went stratospheric.

Modal helped TypeSafe train and deploy a new class of frontier model, then scale to over one trillion tokens in three days while meeting their requirements for ultra-low-latency inference without managing their own infrastructure.

A new class of frontier model

Jev isn’t a large language model. It’s a System One Model–as in fast, intuitive thinking. Jev achieves similar levels of intelligence on System One tasks when compared to frontier LLMs, but focuses on decisions rather than generalized intelligence.

To bring Jev to life, TypeSafe built a custom stack entirely focused on automation: with a new model architecture, parallel sampler for maximum efficiency, and training method they call Reinforcement Learning for Calibrated Decisions (RLCD).

Unlike LLMs, Jev isn’t autoregressive and it doesn’t generate text. Instead, it takes state and questions as input, and outputs typed, structured values with probabilities. For example, take a customer support ticket as input:

The user defines the choice primitive, which directs Jev to choose an option from a list. Given the options sales, technical, or billing , Jev returns outputs and probabilities:

Jev doesn’t generate outputs sequentially, but in parallel–producing intelligent, low-latency responses suitable for workloads like large-scale automation, fast browser use, and generative user interfaces.

Scaling exponentially over a weekend

Jev launched publicly on a Thursday. One week later, it topped the leaderboard on OpenRouter for contexts between 1K-10K tokens, handling 15.1% of requests. This is just a slice of the demand being handled by the TypeSafe team, which they now measure “in the trillions” of tokens.

Jev OpenRouter share.

To scale their systems while serving Jev at the price-performance frontier, the TypeSafe team needed full control of their stack–from networking and routing to autoscaling for unpredictable load. Most importantly, they needed that control without the headache of managing their own infrastructure.

After all, inference is a (very) big thing, but it’s not everything.

Building on Modal

Modal’s infrastructure is uniquely suited for model launches like Jev: bursty, high-concurrency workloads where every millisecond counts. Instead of guessing at capacity needs or overpaying for a compute reservation, AI labs can use Modal’s serverless platform to train, serve, and scale any type of model: LLMs, System One Models, world models, and more.

“Modal served 1 trillion tokens for Jev within three days of launch. Their team was proactive and responsive as we scaled to meet unprecedented demand, and we could stay focused on shipping instead of managing infrastructure.”
Erik Gafni CTO

Jev's launch is one example of what the Modal platform was built for. Bring any model, any serving code, and scale to trillions of tokens per day.

来源:Modal 官方工程博客 · modal.com