Ollaya points at a missing layer for decision models

Ollaya points at a missing layer for decision models

4 min read

Ollaya is pitched as Ollama for open-source, Jev-style decision models, which is less interesting as a single tool than as a signal: decision models need the same boring packaging, local running, and evaluation habits LLMs finally got.

TL;DR: Ollaya’s useful idea is not “another model runner,” it is that decision models need a local, repeatable operating layer before builders can trust them inside real workflows.

What is Ollaya actually suggesting?

The primary source here is the Hacker News submission titled “Ollaya – Ollama for open-source, Jev-style decision models.” That title is doing most of the work. It frames Ollaya as an Ollama-like layer, but for open-source decision models rather than general chat models.

That is a useful framing, even with limited public detail in the submission.

Ollama mattered because it made local model use feel boring. Pull a model. Run it. Swap it. Script against it. The big shift was not raw capability. It was packaging. Developers could stop treating local inference like a science project.

If Ollaya is aiming at that same slot for Jev-style decision models, the important question is not whether it beats a frontier LLM on broad reasoning. It is whether it makes decision behavior easy to install, inspect, compare, and repeat.

That is a different problem from chat. A chatbot can be fuzzy and still useful. A decision model usually sits inside a loop: approve, reject, route, bid, rank, escalate, wait, retry. The output often changes state in another system. That makes drift, hidden assumptions, and weak evals much more expensive.

Why do decision models need different plumbing?

Most AI tooling still assumes the model is producing text for a human. Decision systems are different. They need inputs that look like state, outputs that look like actions, and tests that look like consequences.

A sales agent deciding which lead to contact next is not just “answering.” A support triage system choosing whether to refund, escalate, or ask for more information is not just “generating.” A crawler deciding whether a page is worth fetching again is not writing prose. These are policies, not paragraphs.

That distinction matters because the interface changes what you measure. With language models, builders often test vibe, helpfulness, correctness, refusal behavior, latency, and cost. With decision models, you also care about stability under small input changes, regret over time, calibration, safe fallback behavior, and whether the model gets stuck in a locally smart but globally dumb pattern.

two contrasting machines, one producing flowing text to a person, the other sending small action tokens into a feedback

This is where an Ollama-like pattern could help. Not by making decision models magical, but by making them swappable. If a builder can run Model A and Model B against the same logged states, compare action choices, replay outcomes, and keep that process local, the category becomes easier to reason about.

The catch: a runner does not solve evaluation. It just makes evaluation possible enough that teams lose their excuses.

Where would I try this first?

I would not start with money movement, hiring, medical decisions, or anything where the blast radius is high. Start with low-risk routing and prioritization, where a bad choice is recoverable and observable.

Good first tests: queue ordering, internal task assignment, content moderation pre-labeling, research source triage, crawler scheduling, or choosing which tool an agent should call next. These are places where a decision policy can improve a workflow without pretending to be an autonomous executive.

The practical move is to collect real historical states first. Then run candidate decision models against those states offline. Compare them to what humans or existing rules did. Look for obvious failure clusters. Only then put the model in shadow mode, where it recommends actions without taking them. If the system still looks useful after that, allow narrow production use with logs, override paths, and rollback.

I like the direction because it shifts attention from chat demos to operational control. But I would be careful with the branding gravity around “Ollama for X.” Ollama succeeded because the developer loop was crisp. Decision models need that, plus a much harder eval loop.

Practitioners should treat Ollaya as a prompt to build a decision-model harness, even if they never adopt the tool. Pick one reversible decision in your workflow. Log the state, available actions, chosen action, and result. Test a local open model against that record. The catch most teams miss: the model is the easy part. The asset is the replayable decision history.