Our 4-billion-parameter open model is #1 on JevBench’s official Composite Score, which weighs intelligence, calibration, cost and speed together, and it outscores Jev itself: 72.5 to 71.5, on the benchmark built around Jev. It also beats every 12B, 26B and 31B model on the board. It is Apache-2.0, on Hugging Face today, runs on stock vLLM, and the current release reads images too.
#1
OFFICIAL COMPOSITE
72.5 vs 71.5
COMPOSITE,
US VS JEV
90.0
CALIBRATION
$0.021
PER 1,000 DECISIONS (EST.)
JevBench is a public leaderboard for decision models. Each model receives a record and a question with a fixed set of answers, and returns a probability for each answer.1 Its official Composite Score ranks 104 open-weight systems. H2O-Lightning-4B v1.1 is first: 72.5, ahead of Jev 1.13.0 (71.5), of the 31B Quyet-1.0-Large (71.4), and of every 12B entry.
The JevBench v1.6.1 Composite Score, official equal weights, read 7 October 2026. Jev 1.13.0 is shown as the board’s unranked reference row.2
A Jev-class model is built to make decisions rather than write text. You give it the state (a ticket, a policy, a conversation, a document) and one or more questions, each with a bounded set of answers. It returns a probability for every answer in a single pass, with no generated explanation. JevBench puts it as “state and a bounded rubric in, a typed answer out.”
Many decisions inside real software have this shape: assigning a ticket priority of P1, P2 or P3, checking whether a claim matches its receipt, or rating urgency on a three-point scale. These calls need an answer you can act on, a confidence you can trust, and low cost and latency, because they run thousands of times a day.
The name comes from Jev, TypeSafe AI's hosted decision API, which defines the genre. The board also uses it as a yardstick: to count as Jev-class, a system has to stay within twice Jev's cost and twice its median latency.3
JevBench scores every system on four axes from 0 to 100: Intelligence (does it pick the right answer), Calibration (do its probabilities mean what they say), Speed and Cost. The composite is the equal-weight harmonic mean of the four, so a weak axis drags the total down hard.
| # | System | Size | Composite | Intel. | Calib. | Speed | Cost | $ / 1k |
|---|---|---|---|---|---|---|---|---|
| 1 | H2O-Lightning-4B v1.1 | 4B | 72.5 | 60.0 | 90.0 | 92.6 | 60.3 | $0.021 |
| – | Jev 1.13.0 (reference, hosted API) | n/a | 71.5 | 63.6 | 90.6 | 91.5 | 54.7 | $0.032 |
| 2 | Quyet-1.0-Large | 31B | 71.4 | 73.4 | 90.0 | 86.9 | 50.5 | $0.045 |
| 3 | torchcast-decision-12b | 12B | 69.9 | 60.5 | 82.9 | 91.7 | 56.4 | $0.028 |
| 4 | Winnow-12B Q8 | 12B | 68.9 | 59.5 | 83.0 | 86.7 | 56.6 | $0.028 |
| 5 | deck-31B | 31B | 68.7 | 73.0 | 82.2 | 86.7 | 49.7 | $0.048 |
| 6 | Cygnet | 12B | 68.6 | 54.8 | 87.0 | 91.8 | 56.4 | $0.028 |
| 7 | Jev-Omni | 12B | 67.7 | 55.5 | 87.0 | 85.4 | 56.1 | $0.029 |
| 8 | Xor 26B-A4B | 26B MoE | 67.4 | 58.7 | 86.1 | 91.0 | 50.7 | $0.044 |
JevBench v1.6.1 Composite Score, official 25:25:25:25 weights, read 7 October 2026. Size is the base model the board lists for each row. Costs marked est. on the board are estimates for self-hosted weights.4
The board's own detail card for our row: 0.65× Jev's cost and 0.34× its median latency, both inside the Jev-class limits, and “official #1”.
On raw Intelligence the 31B models are clearly ahead, at 73 against our 60. In production, though, decision systems also need honest probabilities, fast responses and a low cost per call, and JevBench scores all four.
Our lead comes from three of the four axes:
The four axes, H2O-Lightning-4B v1.1 (solid) against the 31B Quyet-1.0-Large (dashed), from the board's compare view. The 31B is well ahead on Intelligence. Calibration is a tie, and the 4B is cheaper and faster.
Capability against cost for the Jev-class systems on the open-weights board, with our bubble selected. Right is cheaper; bubble size follows the composite.
The result comes from a disciplined release process rather than one technique. In brief:
How a candidate becomes a release. Candidates are compared on the held-out lockbox under rules written down in advance, and the gaps we find feed the next round.
The model is at huggingface.co/h2oai/h2o-lightning-4b under Apache-2.0. It runs on unmodified vLLM 0.30.0 with a small standard-library Python shim in front of it, both described on the card. The commands below are the card's own. In a fresh virtual environment:
pip install "vllm==0.30.0+cu129" "torchcodec==0.16.0+cu129" --extra-index-url https://wheels.vllm.ai/0.30.0/cu129 --extra-index-url https://download.pytorch.org/whl/cu129
hf download h2oai/h2o-lightning-4b --revision v1.2.1 --local-dir h2o-lightning-4b
Serve the model, then start the shim on the same machine:
vllm serve ./h2o-lightning-4b --served-model-name h2oai/h2o-lightning-4b --host 127.0.0.1 --port 8000 \
--max-model-len 40960 --gpu-memory-utilization 0.90 \
--limit-mm-per-prompt '{"image": 4, "video": 0}' --mm-processor-kwargs '{"max_pixels": 1605632}'
python3 h2o-lightning-4b/h2o_lightning_shim.py --config h2o-lightning-4b/serve_config.json --vllm http://127.0.0.1:8000 --port 8741
Or run bash h2o-lightning-4b/serve.sh, which does both. Warm up until this returns HTTP 200:
curl -s http://127.0.0.1:8741/v1/systemone -H 'Content-Type: application/json' -d '{"state": "warm-up",
"questions": {"decision": {"type": "noul", "instructions": "Is this a warm-up?",
"criteria": {"true": "yes", "false": "no"}}}}'
Then send real decisions to POST /v1/systemone. One request can carry several questions about the same state. This is the card's example:
{"state": "Ticket 4471: since 09:10 the checkout page returns HTTP 500 for every customer; no orders have gone through in 40 minutes. Severity guide: P1 = a revenue path fully down; P2 = degraded but working; P3 = cosmetic.",
"questions": {
"priority": {"type": "choice", "instructions": "Which priority does the severity guide assign?",
"criteria": {"p1": "Revenue path fully down", "p2": "Degraded but working", "p3": "Cosmetic"}},
"all_users": {"type": "noul", "instructions": "The problem affects every customer.",
"criteria": {"true": "Affects everyone", "false": "Affects only some"}},
"urgency": {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["Low", "Medium", "High"]}}}
Every answer comes back as probabilities: P1 at 0.944, “affects everyone” at 0.929, and a full distribution over Low, Medium and High, in the card's own example. There are three question types: choice picks one of a set of named options, noul gives the probability that a statement is true, and score gives a level on a stated scale.
The current public release, v1.2.1, adds image input. Put a photo, a screenshot, a scanned document or a chart in the request as a data URI, and the same three question types work over it: does the receipt total match the claim, how many item lines are there, how legible is it. Up to four images go in one request. A request with no image takes exactly the text path that was measured on the board.6
The model card on Hugging Face. Apache-2.0, built on Qwen/Qwen3.5-4B, served with unmodified vLLM and an open shim.
The board can also display hosted API offerings, unranked and for context. Some of them score higher on the same formula (Sage 1.3.0 at 74.0, Liquid AI’s d1 at 73.0), but they are not part of this ranking and have a board of their own.7 On the Capability-only view, which averages Intelligence and Calibration and leaves cost and speed out, H2O-Lightning-4B is #3 among open models, behind two 31B models.
H2O-Lightning-4B shows how far a small, well-calibrated model can go when cost and speed count. Larger H2O-Lightning models are on the way, aimed at the Intelligence gap without giving up what makes this one fast and cheap. We’ll share them when they’re ready.
In the meantime, you can download the 4B, run it on a single GPU, and follow the leaderboard at benchmarkheaven.com/jev-models.
Our thanks to Florian Standhartinger, who created JevBench and runs Benchmark Heaven. An independent benchmark with sealed items, re-measured rows and public method notes makes results like this one credible, and gives everyone building Jev-class models a fair way to compare and improve. Thanks as well to the Qwen team for the Apache-2.0 Qwen3.5-4B base model, and to the vLLM project, which H2O-Lightning runs on unmodified.
H2O-Lightning sits alongside the rest of the H2O.ai platform:
See h2o.ai for the full platform.
Notes