SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does

SWE-Flux: The Benchmark That Asks If Coding Models Can Predict What Code Actually Does

6 min read

A new repository-level benchmark tests whether LLMs can reason about runtime behavior, not just read code. The best model scored 37 percent. Here is what that gap means for anyone shipping coding agents in production.

TL;DR: A new benchmark called SWE-Flux tests whether LLMs can predict how real code behaves when it runs, and the best model got 37 percent right, which means your coding agent can read a repo far better than it can tell you what that repo will actually do.

The primary source here is “Can LLMs Reason About Runtime Behavior? A Repository-Level Dynamic Benchmark,” which introduces SWE-Flux, posted to arXiv under both cs.AI and cs.CL. It does something most coding benchmarks skip. It asks the model not “what does this code mean” but “what happens when it runs.”

That distinction is the whole story.

What is SWE-Flux actually testing?

Most repository-level benchmarks measure static understanding. Can the model explain a function, find where a symbol is defined, summarize a module. Useful, but it is reading comprehension. SWE-Flux goes after execution reasoning: given real code and a real test, can the model predict the control flow, the loop iterations, the program state at a given point, how data moves, which exceptions fire, and which invariants hold.

The benchmark contains 480 execution-grounded instances across 12 real Python repositories. The questions span single-test and multi-test scenarios covering control flow, loops, program state, dataflow, exceptions, and program invariants.

The clever part is how the answers get made. The authors do not write gold answers by hand, and they do not use an LLM to judge correctness, which is the shortcut that quietly poisons a lot of coding evals. Instead they instrument the test executions and harvest the ground truth directly from what the code did when it ran. The oracle is the interpreter, not a person’s guess and not another model’s opinion.

a static blueprint of a machine on one side and the same machine actually running with moving parts on the other, showin

That matters because runtime is not up for debate. A loop ran 14 times or it did not. An exception was raised or it was not. When your gold answers come from instrumented execution, you remove the two biggest sources of benchmark rot: human labeling error and LLM-as-judge bias. This is the methodological point worth stealing even if you never touch SWE-Flux itself.

Why did the best model only hit 37 percent?

Five LLMs were evaluated. The best one reached 37 percent accuracy. Read that again in the context of how good these same models look on SWE-bench-style task completion, where the leaderboard numbers keep climbing. On predicting runtime behavior, they fall off a cliff.

The failure pattern is specific and it tells you something real. Models did better on localized behavior: invariants, control flow inside a single procedure, exceptions, simple loops. They struggled on dataflow, inter-procedural execution, precise state reasoning, and suite-level aggregation.

Notice the shape of that. The things models handle are the things you can reason about by looking at one chunk of code in isolation. The things they fail at all require tracking state across function boundaries and over time. Inter-procedural execution means following the call chain. Precise state reasoning means holding a mental model of what every variable actually contains after N steps. Suite-level aggregation means reasoning about many tests at once.

a single glowing node that a viewer can trace easily versus a tangled web of connected nodes spreading across many regio

This is the same weakness that shows up when you watch a coding agent work on a large codebase. It writes a plausible patch that looks right locally, then breaks something three files away because it never actually simulated what the change would do at runtime. The model pattern-matches to “code that looks correct” rather than executing a mental model of the program. SWE-Flux puts a number on that intuition, and the number is not flattering.

Does this contradict the “AI writes code” story?

No, and this is where I want to cut against both the hype and the backlash. Writing code and predicting code behavior are different skills, even for humans. Plenty of good engineers write working code without being able to trace exact runtime state in their head, because they lean on the interpreter, tests, and a debugger to close the gap.

That is the key operator lesson. The models are bad at simulating execution in their heads, so stop making them do it in their heads. Give them the runtime. An agent that can run the code, read the actual stack trace, inspect the real variable values, and iterate will beat an agent asked to reason about behavior cold. SWE-Flux measures the cold-reasoning ceiling, and that ceiling is low. Your production setup should route around it.

I would push back gently on one thing. Five models is a small panel, and the abstract does not name them here, so I am not going to tell you which frontier model tops out at 37 percent or how a reasoning-heavy model with more test-time compute would fare on the same questions. It is plausible that models built to spend more tokens tracing execution would do meaningfully better. The paper does not give that breakdown in the material I have, so I am treating the 37 percent as a snapshot of the models tested, not a universal law.

One more useful bit: the authors show the oracle-harvesting pipeline can generate fresh benchmark variants through input perturbation, successfully harvesting valid variants for nearly 90 percent of selected instances, and those variants are substantially harder for the models. That is a real defense against benchmark contamination. When your eval can regenerate itself, memorizing the test set stops working.

What should a builder do with this?

The practical read is about architecture, not despair. If you are shipping a coding agent, the SWE-Flux result says your agent’s internal belief about what code does is unreliable, especially across function boundaries and over multiple steps. So do not trust it. Build the loop that grounds every claim in execution.

Concretely: wire your agent to run tests before and after edits, feed it the real stack traces instead of asking it to predict them, and give it a way to print or inspect intermediate state rather than reasoning about state in prose. Treat any agent output that asserts runtime behavior without having run the code as a hypothesis, not a fact. If you are running evals on your own agent, borrow the harvesting idea: derive ground truth from instrumented execution instead of an LLM judge, and perturb inputs to keep the model from memorizing. The catch most readers will miss is that a rising SWE-bench score does not fix this. Task-completion benchmarks reward getting to a passing test, often through iteration and tooling. They do not measure whether the model understood why it passed. Those are different competencies, and if your product depends on the model reasoning about behavior it cannot observe, this 37 percent is the number you are actually building on top of.