A 3D character zipping through a sunny city on a green skateboard with green motion trails and floating completed-task notifications - asi1-mini moving fast through high-volume work.
New Model

Meet asi1-mini: Your AI in a Hurry

The fastest, lightest model in the ASI:One family. Built for the workloads where speed is the feature - real-time chat, voice, classification, and the kind of bounded tasks you run a million times a day.

May 6, 2026•5 min read

Some questions deserve thirty seconds of careful thought.

"What category does this support ticket belong to?" is not one of them. Neither is "what should the next word be in this sentence," or "is this message a complaint or a compliment," or "given these three options, which one should I take." Those questions need to be answered fast, cheaply, and a few hundred times a minute.

That is the work asi1-mini was built for.

What asi1-mini Is

asi1-mini is the fastest, lightest model in the ASI:One family. It shares the same API surface as the rest of the family and the same full 200,000-token context window - so you can give it the same inputs you would give the bigger models. What you trade for speed is reasoning depth and response size, not the ability to read context.

  • Built for low-latency workloads Tuned for fast first-token times and short, focused responses. The right pick for anywhere a user is waiting on the screen.
  • 200,000 tokens of context Pass it the full document, the entire ticket history, the long system prompt - it can read it. The speed comes from how it reasons and responds, not from a smaller window.
  • Same tool calling, same APIs Full tool-calling support. OpenAI SDK compatible. Streaming. Same Chat Completions and Responses APIs as the rest of the family.
  • Lowest cost in the family Designed for high volume. The workloads you run thousands or millions of times a day stay economical.

Where asi1-mini Shines

The pattern: bounded tasks, high volume, latency matters. If your workload looks like this, asi1-mini is probably the right call.

  • Real-time chat A user typing on the other end of a conversation. The faster the first token, the better the experience. asi1-mini keeps the loop tight.
  • Voice assistants Voice is the most latency-sensitive interface there is. Half a second of waiting is half a second too long. asi1-mini is built for the voice loop.
  • Classification and routing Triage support tickets. Label incoming events. Detect intent. Decide which queue a message belongs in. Short input, short output, well-defined task - asi1-mini handles it.
  • Autocomplete and inline suggestions Code suggestions, email drafting, search-as-you-type. Workloads where the user is mid-thought and the AI has to keep up.
  • Simple tool calls When the model is choosing between a small set of well-defined actions, you do not need deep reasoning - you need a fast, confident answer. asi1-mini delivers it.
  • High-volume background work Bulk summarization, batch enrichment, automated tagging. The kind of work where the cost-per-call decides whether the project is feasible at all.

Same Family, Different Tradeoff

The three ASI:One models share everything that makes the family useful: the same API, the same tool calling, the same OpenAI SDK compatibility, the same 200K context window. They differ in reasoning depth, response size, and speed.

That means you can mix them in the same application. Use asi1-mini for the fast paths, asi1 for the general ones, and asi1-ultra for the deep work. Or cascade them: try asi1-mini first, fall back to asi1 when the smaller model is uncertain. Switching between models is a one-line change.

A useful pattern

Run asi1-mini on the request first. If it is confident, ship the answer. If it is not, retry with asi1 or asi1-ultra. Most of your traffic is bounded and confident; you only pay for depth when you actually need it.

How to Use It

For developers, asi1-mini is live in the API today. Same key. Same request format. Just change the model field:

Switching is one line

Change "model": "asi1" to "model": "asi1-mini" and you are done. Same context window. Same tools. More speed.

See the asi1-mini developer docs for specifications, code examples, and a migration walkthrough. For a side-by-side comparison and a decision tree, see Model Selection.

When Not to Use It

asi1-mini is the wrong pick when depth is what you need. If your task involves long agentic flows, multi-hop research, code review of a real codebase, or any work where the answer requires connecting evidence across many steps, use asi1-ultra. If your task is general-purpose and varies in complexity, use asi1.

A useful tell: if you find yourself wishing the response were longer or more thorough, you have outgrown asi1-mini for that workload. Move it to asi1.

What This Unlocks

Speed and cost change what is possible to build. Real-time AI features that were too expensive at scale - chat that does not stall, voice that does not lag, classification on every event in your pipeline - are now in reach.

asi1-mini is the model for the work that has to happen fast and often.

Keep Reading

Build something fast.

Switch your model field to asi1-mini and ship the real-time AI feature you have been holding back. Or pick up the developer docs and start prototyping.