Perplexity Just Launched a Decisions API — An AI That Answers in Probabilities, Not Words
Perplexity just did something unusual: it launched an AI model that doesn't generate text. On October 1, 2026, the company announced the Decisions API, powered by a new model called pplx-decider-v1-27b — a system that takes in text and images and answers in probabilities instead of words.
Think of it as the opposite of a chatbot. Instead of writing a paragraph, you hand the model a fixed set of possible answers — route this ticket to billing or support, classify this intent, pick the right tool for the job — and it returns calibrated probabilities over those options in a single non-autoregressive decision step. No token-by-token generation, no parsing free-form output, no surprises.
The details
- The model: pplx-decider-v1-27b is a multimodal decision model fine-tuned from Qwen3.8-27B, with roughly 26 billion parameters. It handles both text and images as input.
- How it works: instead of generating free-form text, it outputs calibrated probabilities over a fixed set of answers in one non-autoregressive decision step — purpose-built for routing tickets, classifying intent, and picking tools.
- Pricing: $0.04 per million input tokens, and $0 for output. Because decisions come back as probabilities rather than generated text, the output side of the bill is effectively zero.
- Performance: Perplexity reports 85.71% overall accuracy across 11 benchmarks.
- Open weights: the model weights are released under the Apache 2.0 license, with a 262,144-
token context window — so teams can inspect it, fine-tune it, or self-host it.
Why it matters
AI agents spend a large share of their compute — and their latency — on decisions that never needed prose in the first place. Every time an agent calls a big language model just to decide which tool to use or which queue a ticket belongs in, it pays for full text generation and then has to parse the answer back into a structured choice. That is slow, expensive, and fragile.
A dedicated decision layer flips that equation. One fast probabilistic call replaces a generate-then-parse round trip, at a fraction of the cost. At $0.04 per million input tokens with free output, high-volume classification and routing workloads suddenly get dramatically cheaper to run. And because the weights are open, developers aren't locked into a single vendor's API to use it.
The takeaway
Perplexity is betting that the next phase of AI isn't about models that talk more — it's about models that decide faster. If agents are going to run at real-world scale, they need a cheap, structured decision layer sitting underneath all that generated text. With open weights, near-zero output cost, and strong benchmark numbers, the Decisions API is a serious entry in that race. Agents just got a decision layer for spare change.

Comments
Post a Comment