Introducing Lake Ultra R1
A decision model for routing user prompts to the right model with subsecond latency.
Over the past decade, advances in AI have uncovered a new set of problems which machines can solve. While frontier models continue to push those boundaries, smaller models offer new combinations of accuracy, latency, and cost.
This creates an opportunity to rethink how intelligence is deployed. Different tasks require different capabilities, and greater computation does not always translate into a better outcome. Making AI economically scalable requires understanding where additional intelligence improves results and where efficient models already meet the demands of the task.
The next stage of AI progress will depend on how effectively we deploy intelligence, alongside how far we advance its capabilities.
Today, we are excited to announce Lake Ultra R1, the first model in the Ultra family.
R1 is a decision model designed to route requests to the right model in subsecond latency.

Figure 1. Task classification accuracy on the held-out benchmark.
On our task classification, R1 delivers performance competitive with GPT 6 Astra, with significantly reduced costs and latency. On our task classification benchmark, R1 achieved 91.3% accuracy, compared to GPT 6 Astra’s 93.1% and outperforming Claude Opus 5.5 at 79.0% and Jev at 75.1%.
R1 is highly specialized to reduce latency and cost, its mean response time is 17.9x faster than GPT 6 Astra and 6.4x faster than Claude Haiku 4.5 on the same request.

Figure 2. Mean end-to-end request latency on the same request set.
Further, R1 reduces per request cost by approximately 1,000x relative to GPT 6 Astra and 580x relative to Claude Opus 5.5, accounting for input, output, and cache charges for our workload configuration. Lake Ultra R1 is priced at $0.033 per million input tokens. Output tokens are always free.

Figure 3. Cost per thousand requests, including the input, output, and cache charges for each tested configuration.

Figure 4. Classification accuracy versus cost per thousand requests.
The training builds on a pretrained model through supervised fine tuning and RLHF. Applications use R1’s probabilities to match each incoming user request to a model with the appropriate capabilities, based on quality requirements, latency targets, and inference budgets.
Beyond the cost of classification, fast routing can reduce the cost/latency of executing requests. In our benchmarks we used the Ultra R1 to direct messages to equally capable smaller open source models which reduced cost per request by 68.6% relative to cached GPT 6 Astra. This marks a significant shift in how models are used as enterprises aim to balance cost with intelligence. This approach allows applications to retain frontier models for demanding work while assigning suitable requests to more efficient models.
For an enterprise application, these capabilities work together. Lighter tasks are assigned to efficient but smaller models, while requests involving more demanding reasoning receive frontier capability. Model selection allows teams to balance quality, response time, and expenditure across the work their systems perform.
R1 classifies requests into eight task categories: classification, light coding tasks, coding, complex reasoning, ambiguous problems, consequential decisions, computer use, and tool/API use. It returns a probability for each category, allowing applications to set their own thresholds and apply stricter thresholds for tasks with greater consequences. The dataset consists of real world user chat prompts across several domains including general purpose coding, reasoning, tool use and classification with English as the dominant language category. Evaluation uses a held out test set which covers the same categories in approximately the same proportions as the full training set and measures how often the model’s highest probability category matches the reference label.
R1 uses a hybrid attention architecture attached to a decision head to produce category probabilities. The backbone contains 64 blocks, arranged as three gated linear attention blocks followed by one full attention block. The linear attention carries information through a fixed recurrent state, with gates controlling how that state incorporates new information. Full attention blocks provide direct access to previous token representations, allowing information from distant parts of the request to be attended to. Finally, the decision head converts the output representation into category scores, normalized to produce a probability distribution over task categories.
We identified two opportunities to reduce memory traffic during inference. In both cases, an intermediate tensor is produced solely to be consumed by the next operation. Keeping these values on GPU registers allows both operations to be executed in a single kernel, eliminating the intermediate write to GPU memory, the subsequent read, and avoiding a new kernel launch.
For linear attention, we implement a fused convolution and SiLU kernel, parallelized across token positions. Each output combines the current and preceding token representations, applies SiLU locally, and writes the result directly to its destination. For the feed forward layers, a fused elementwise kernel computes the activation and applies the gate leaving the projections unchanged. These optimizations apply for the whole backbone and together reduce forward pass time by 8% against the unfused implementation.
These optimizations reduce the computational overhead of routing.
Developers can integrate Lake Ultra R1 into existing agents, application workflows, and inference gateways. Each request produces structured predictions that an application uses to select the model responsible for the response, according to its quality requirements, latency targets, and inference budget.