跳到正文
Liquid AI 模型与工程博客·· 19 小时前精选AI 评分69

Liquid AI 发布 d1 决策模型并新增图像输入能力

Introducing d1: The most capable decision model, now with vision

AI 导读

Liquid AI 发布 d1 决策模型,新增文本与图像输入,可通过 console.liquid.ai 和 d1 Playground 使用。

推荐理由

原文给出 d1 在六类真实应用中对 GPT-6.1 Sol 和 Claude Opus 5.5 的成本与速度对比,以及按输入 token 计费规则,便于评估是否替换现有 LLM 调用。

正文 · 原文

Today, we introduce d1, our first decision model, now supporting both text and images. With last week's experimental release, d1 became the first model to rival Jev on text decisions, and it is now the first to extend these capabilities to images. You can try it today atconsole.liquid.ai and in the d1 Playground.

We tested d1 against GPT-6.1 Sol and Claude Opus 5.5 on six real applications, from filtering support tickets to inspecting circuit boards. d1 matches or beats GPT-6.1 Sol on four of them. It costs 19x to 200x less than both models and answers significantly faster on every task.

How the d1 decision model works

Decision models answer questions about a situation with a probability for each possible answer. The d1 decision model takes unstructured data (e.g., text, images, or both) and one or more questions as input, reads them in one forward pass, and returns the probabilities, without generating any tokens. A text decision takes 200 to 300 ms, fast enough for real-time applications.

d1 answers three types of questions:

  • Noul: a yes/no question, answered with a probability between 0 and 1.
  • Choice: pick one label among many, answered with a probability per label.
  • Score: a position on a scale, weighted by the probability of each level.

One request can ask several questions about the same state, saving input tokens.

d1 in action

Decision models can replace expensive calls to language models when the answer is a structured decision. They also enable a variety of new use cases, some shown in this section. Every demo below runs live in the d1 Playground.

Visual inspection. This industrial application reviews parts from four production lines that pass under a camera: circuit boards, candles, cashews, and chewing gum (public VisA dataset). d1 sorts good and defective parts with 85-97% accuracy. The most interesting part is that the model was never trained for this. Thanks to its excellent generalizability, it understands the task from a short description.

Text applications. Five applications use d1 for every decision:

  • Smart Filter: d1 becomes a function in a SQL query, WHERE d1(ticket, 'the customer wants to cancel). It answers yes or no for each of 150 support tickets.
  • Code Search: d1 browses the Hugging Face transformers repository (6,511 files) one folder at a time, down to the function that answers a question.
  • Smart Folders: d1 files each new document into a folder, then a subfolder. It files search questions the same way, so keyword search only looks in one subfolder.
  • Web Agent: d1 operates a flight-search website from a one-sentence goal. At each step, it picks the next action among everything the page allows.
  • Context Compaction: d1 reads each tool output in a coding agent's session and keeps, trims, or drops it for the next task. It removes 52% of the tokens and keeps every output the task needs.

We adapted four of these applications from the following open-source projects: pg-jev, jevgrep, jev-ultrafast, and fast-jev-compaction.

Games. d1 plays eight classic games live and picks every move. Vision helps it in two ways:

  • Better decisions: Tetris can be fully described in text, but adding the screen raises d1's score from 70 to 81 cleared lines.
  • Simpler integration: In Wordle, d1 reads the board directly from a screenshot. Developers don't need to write a text version of the game, and d1 still solved 12 of 12 games in 3.8 guesses on average.

d1 also handles purely visual tasks. In Quick, Draw!, it guesses what a player's doodle shows among 62 words, and recognizes 5.2 of 6 drawings (random guessing gets 0.6).

Availability and pricing

Start building today with the d1 decision model, available on the Liquid AI API as d1. Create an API key at console.liquid.ai (Dashboard > API Keys).

To use vision, you can simply send images as base64 data URLs in images:

import base64, os, requests

image = base64.b64encode(open("board.jpg", "rb").read()).decode()

response = requests.post(
    "https://api.liquid.ai/decisions/v1/systemone",
    headers={"Authorization": f"Bearer {os.environ['LIQUID_API_KEY']}"},
    json={
        "model": "d1",
        "images": [f"data:image/jpeg;base64,{image}"],
        "state": "Camera image of a circuit board on the production line.",
        "questions": 
            {"defect": {
                "type": "noul", 
                "instructions": "Does this circuit board have a defect?"
             }
         },
    },
)

print(response.json()["answers"]["defect"]["noul"])

The full API reference is in the decision models documentation.

d1 is billed on input tokens only, with no output tokens. Images are counted as input tokens at the same rate as text: 1.5 tokens per 32×32-pixel patch, so a 1024×1024 image costs 1,536 tokens. Each question is billed as its own prompt, including its text and all images.

d1 is also available through Vercel and OpenRouter, with text only for now. Vision is coming to both soon.

This is an exciting time for creativity and exploration in AI, and d1 is only the beginning. We're continuing our work on decision models across a range of sizes, with new features, higher decision quality, and lower latency, and we plan to release open weights for upcoming models on Hugging Face soon.

Liquid AI logoTry on Liquid ConsoleLiquid AI logoRead our docs

Citation

For citations, please use the following reference or BibTeX:

Liquid AI, "Introducing d1: The most capable decision model, now with vision", Liquid AI Blog, Oct 2026.

Methodology. We ran each application once per model on October 5, 2026, with the d1 Playground's comparison script. GPT-6.1 Sol and Claude Opus 5.5 get each request as one chat message and answer in JSON, at their default reasoning setting. Where many questions share one input, such as the filter's tickets or the folders' passages, the chat models answer them in batches. Costs use list prices, without prompt-cache discounts. d1's costs use $0.04 per million input tokens. Time is per run, with up to 8 requests in flight. A Smart Filter run is one query over 150 tickets. A Smart Folders run files 105 passages, and its cost is per 1,000 passages. The filter's quality is its F1 score against hand labels, which counts both missed and wrong matches. The other text applications count the goals, questions, passages or needed outputs handled correctly. We wrote six of the 15 code questions and two of the four compaction sessions after d1's pipeline was set. In Visual Inspection, every model sees a good part from the same line next to the part to inspect.

来源:Liquid AI 模型与工程博客 · liquid.ai