decorator decorator
Return to page
Search
decorator decorator
Uncategorized

H2O-Lightning-4B: #1 on JevBench, and Ahead of Jev Itself

Published: October 07, 2026 min read Written by: Jonathan McKinney
decorator

Our 4-billion-parameter open model is #1 on JevBench’s official Composite Score, which weighs intelligence, calibration, cost and speed together, and it outscores Jev itself: 72.5 to 71.5, on the benchmark built around Jev. It also beats every 12B, 26B and 31B model on the board. It is Apache-2.0, on Hugging Face today, runs on stock vLLM, and the current release reads images too.

#1

OFFICIAL COMPOSITE

72.5 vs 71.5

COMPOSITE,
US VS JEV

90.0

CALIBRATION

$0.021

PER 1,000 DECISIONS (EST.)

JevBench is a public leaderboard for decision models. Each model receives a record and a question with a fixed set of answers, and returns a probability for each answer.1 Its official Composite Score ranks 104 open-weight systems. H2O-Lightning-4B v1.1 is first: 72.5, ahead of Jev 1.13.0 (71.5), of the 31B Quyet-1.0-Large (71.4), and of every 12B entry.

The JevBench v1.6.1 Composite Score, official equal weights, read 7 October 2026. Jev 1.13.0 is shown as the board’s unranked reference row.2 The JevBench v1.6.1 Composite Score, official equal weights, read 7 October 2026. Jev 1.13.0 is shown as the board’s unranked reference row.2
The JevBench v1.6.1 Composite Score, official equal weights, read 7 October 2026. Jev 1.13.0 is shown as the board’s unranked reference row.2

 

What a Jev-Class Model Is

A Jev-class model is built to make decisions rather than write text. You give it the state (a ticket, a policy, a conversation, a document) and one or more questions, each with a bounded set of answers. It returns a probability for every answer in a single pass, with no generated explanation. JevBench puts it as “state and a bounded rubric in, a typed answer out.”

Many decisions inside real software have this shape: assigning a ticket priority of P1, P2 or P3, checking whether a claim matches its receipt, or rating urgency on a three-point scale. These calls need an answer you can act on, a confidence you can trust, and low cost and latency, because they run thousands of times a day.

The name comes from Jev, TypeSafe AI's hosted decision API, which defines the genre. The board also uses it as a yardstick: to count as Jev-class, a system has to stay within twice Jev's cost and twice its median latency.3

 

The Results

JevBench scores every system on four axes from 0 to 100: Intelligence (does it pick the right answer), Calibration (do its probabilities mean what they say), Speed and Cost. The composite is the equal-weight harmonic mean of the four, so a weak axis drags the total down hard.

#SystemSizeCompositeIntel.Calib.SpeedCost$ / 1k
1H2O-Lightning-4B v1.14B72.560.090.092.660.3$0.021
–Jev 1.13.0 (reference, hosted API)n/a71.563.690.691.554.7$0.032
2Quyet-1.0-Large31B71.473.490.086.950.5$0.045
3torchcast-decision-12b12B69.960.582.991.756.4$0.028
4Winnow-12B Q812B68.959.583.086.756.6$0.028
5deck-31B31B68.773.082.286.749.7$0.048
6Cygnet12B68.654.887.091.856.4$0.028
7Jev-Omni12B67.755.587.085.456.1$0.029
8Xor 26B-A4B26B MoE67.458.786.191.050.7$0.044

 

JevBench v1.6.1 Composite Score, official 25:25:25:25 weights, read 7 October 2026. Size is the base model the board lists for each row. Costs marked est. on the board are estimates for self-hosted weights.4

JevBench JevBench
The board's own detail card for our row: 0.65× Jev's cost and 0.34× its median latency, both inside the Jev-class limits, and “official #1”.

 

Why a 4B Can Win

On raw Intelligence the 31B models are clearly ahead, at 73 against our 60. In production, though, decision systems also need honest probabilities, fast responses and a low cost per call, and JevBench scores all four.

  • When cost and speed count, being nearly eight times smaller works in our favor.

Our lead comes from three of the four axes:

  • Calibration. 90.0, level with Quyet-1.0-Large and within a point of Jev. When the model reports 90%, it is right about 90% of the time, which matters when you set thresholds or escalation rules on its output.
  • Cost. About $0.021 per 1,000 decisions on the board's estimate, against $0.028 for the 12B rows, $0.045–0.048 for the 31B rows, and $0.032 for Jev.
  • Speed. A Speed score of 92.6. The board measured a raw median of 29 ms per decision and, as it does for every self-hosted row, adds a fixed allowance before scoring (0.21 s on that basis, about a third of Jev's v1.5 reference latency).5
The four axes, H2O-Lightning-4B v1.1 (solid) against the 31B Quyet-1.0-Large (dashed), from the board's compare view. The 31B is well ahead on Intelligence. Calibration is a tie, and the 4B is cheaper and faster. The four axes, H2O-Lightning-4B v1.1 (solid) against the 31B Quyet-1.0-Large (dashed), from the board's compare view. The 31B is well ahead on Intelligence. Calibration is a tie, and the 4B is cheaper and faster.
The four axes, H2O-Lightning-4B v1.1 (solid) against the 31B Quyet-1.0-Large (dashed), from the board's compare view. The 31B is well ahead on Intelligence. Calibration is a tie, and the 4B is cheaper and faster.
Capability against cost for the Jev-class systems on the open-weights board, with our bubble selected. Right is cheaper; bubble size follows the composite. Capability against cost for the Jev-class systems on the open-weights board, with our bubble selected. Right is cheaper; bubble size follows the composite.
Capability against cost for the Jev-class systems on the open-weights board, with our bubble selected. Right is cheaper; bubble size follows the composite.
 

 

How We Got There

The result comes from a disciplined release process rather than one technique. In brief:

gap hunting: weak spots found → new, verified data → next round gap hunting: weak spots found → new, verified data → next round
How a candidate becomes a release. Candidates are compared on the held-out lockbox under rules written down in advance, and the gaps we find feed the next round.
 
 
  • A held-out lockbox. A separate set of independently written evaluation items, never trained on, is used only to compare candidate models before a release. Current rules also veto any candidate that is clearly worse there, however good it looks elsewhere.
  • Rules written down first. The ship / no-ship criteria are fixed before results are seen, and we do not tune them afterwards.
  • No benchmark test items in training. JevBench's public items and the test items of every benchmark we use never enter training: we dedupe all training data against every evaluation set, exact and near-duplicate, so scores are not inflated and the model is not overconfident.
  • Gap hunting. We compare ourselves with other systems by topic, language and question type, run adversarial and red-team sets, and fill the gaps we find with new, verified data.
  • Data quality first. Real, licence-clean data first; labels from human annotation or from agreement among several independent strong models; audits that catch and fix flawed data before it reaches training.
  • A curriculum with replay. Training moves from broad to narrow to hard, and every earlier skill keeps a share of each later stage, so nothing learned early is forgotten.
  • Calibration as a first-class goal. Confidence is measured on held-out data, so a stated 70% means right about 70% of the time.
  • Built for cost and speed. One forward pass per decision, no text generation, stock vLLM, and a 4B model.
     

Run It

The model is at huggingface.co/h2oai/h2o-lightning-4b under Apache-2.0. It runs on unmodified vLLM 0.30.0 with a small standard-library Python shim in front of it, both described on the card. The commands below are the card's own. In a fresh virtual environment:

pip install "vllm==0.30.0+cu129" "torchcodec==0.16.0+cu129" --extra-index-url https://wheels.vllm.ai/0.30.0/cu129 --extra-index-url https://download.pytorch.org/whl/cu129
hf download h2oai/h2o-lightning-4b --revision v1.2.1 --local-dir h2o-lightning-4b

Serve the model, then start the shim on the same machine:

vllm serve ./h2o-lightning-4b --served-model-name h2oai/h2o-lightning-4b --host 127.0.0.1 --port 8000 \
  --max-model-len 40960 --gpu-memory-utilization 0.90 \
  --limit-mm-per-prompt '{"image": 4, "video": 0}' --mm-processor-kwargs '{"max_pixels": 1605632}'

python3 h2o-lightning-4b/h2o_lightning_shim.py --config h2o-lightning-4b/serve_config.json --vllm http://127.0.0.1:8000 --port 8741

Or run bash h2o-lightning-4b/serve.sh, which does both. Warm up until this returns HTTP 200:

curl -s http://127.0.0.1:8741/v1/systemone -H 'Content-Type: application/json' -d '{"state": "warm-up",
  "questions": {"decision": {"type": "noul", "instructions": "Is this a warm-up?",
  "criteria": {"true": "yes", "false": "no"}}}}'

Then send real decisions to POST /v1/systemone. One request can carry several questions about the same state. This is the card's example:

{"state": "Ticket 4471: since 09:10 the checkout page returns HTTP 500 for every customer; no orders have gone through in 40 minutes. Severity guide: P1 = a revenue path fully down; P2 = degraded but working; P3 = cosmetic.",
 "questions": {
   "priority":  {"type": "choice", "instructions": "Which priority does the severity guide assign?",
                 "criteria": {"p1": "Revenue path fully down", "p2": "Degraded but working", "p3": "Cosmetic"}},
   "all_users": {"type": "noul", "instructions": "The problem affects every customer.",
                 "criteria": {"true": "Affects everyone", "false": "Affects only some"}},
   "urgency":   {"type": "score", "instructions": "How urgent is this ticket?", "criteria": ["Low", "Medium", "High"]}}}

Every answer comes back as probabilities: P1 at 0.944, “affects everyone” at 0.929, and a full distribution over Low, Medium and High, in the card's own example. There are three question types: choice picks one of a set of named options, noul gives the probability that a statement is true, and score gives a level on a stated scale.

 

It Reads Images Too

The current public release, v1.2.1, adds image input. Put a photo, a screenshot, a scanned document or a chart in the request as a data URI, and the same three question types work over it: does the receipt total match the claim, how many item lines are there, how legible is it. Up to four images go in one request. A request with no image takes exactly the text path that was measured on the board.6

The model card on Hugging Face. Apache-2.0, built on Qwen/Qwen3.5-4B, served with unmodified vLLM and an open shim. The model card on Hugging Face. Apache-2.0, built on Qwen/Qwen3.5-4B, served with unmodified vLLM and an open shim.
The model card on Hugging Face. Apache-2.0, built on Qwen/Qwen3.5-4B, served with unmodified vLLM and an open shim.
 

 

About the API Offerings and the Capability Column

The board can also display hosted API offerings, unranked and for context. Some of them score higher on the same formula (Sage 1.3.0 at 74.0, Liquid AI’s d1 at 73.0), but they are not part of this ranking and have a board of their own.7 On the Capability-only view, which averages Intelligence and Calibration and leaves cost and speed out, H2O-Lightning-4B is #3 among open models, behind two 31B models.

 

What’s Next

H2O-Lightning-4B shows how far a small, well-calibrated model can go when cost and speed count. Larger H2O-Lightning models are on the way, aimed at the Intelligence gap without giving up what makes this one fast and cheap. We’ll share them when they’re ready.

In the meantime, you can download the 4B, run it on a single GPU, and follow the leaderboard at benchmarkheaven.com/jev-models.

 

Acknowledgements

Our thanks to Florian Standhartinger, who created JevBench and runs Benchmark Heaven. An independent benchmark with sealed items, re-measured rows and public method notes makes results like this one credible, and gives everyone building Jev-class models a fair way to compare and improve. Thanks as well to the Qwen team for the Apache-2.0 Qwen3.5-4B base model, and to the vLLM project, which H2O-Lightning runs on unmodified.

 

More from H2O.ai

H2O-Lightning sits alongside the rest of the H2O.ai platform:

  • Enterprise h2oGPTe: enterprise GenAI and AI agents, with multi-model support, cost controls and app integrations, deployable in your own environment.
  • H2O AI Super Agent: autonomous agents that plan and carry out multi-step work across documents, data and tools.
  • Cascade (aka Enterprise LLM Studio): no-code training and fine-tuning of LLMs and small language models.
  • H2O Danube3: lightweight, offline-capable open-weight small language models.
  • H2OVL Mississippi: open vision-language models for OCR and document AI.
  • H2O Driverless AI: AutoML with automatic feature engineering and explainability.
  • H2O Hydrogen Torch: no-code deep learning for image, text and time-series models.
  • H2O-3: the Apache-licensed open-source ML platform for Python and R, with a commercially supported Secure edition.
  • H2O MLOps: deploy, monitor and manage models from training to production.
  • Eval Studio and H2O MRM: automated testing, human-calibrated evaluation and risk monitoring for AI systems.

See h2o.ai for the full platform.

Notes

  1. JevBench is run by Benchmark Heaven, independently of H2O.ai. Its own description: “Benchmark Heaven’s benchmark for Jev-class decision models: state and a bounded rubric in, a typed answer out. It measures Intelligence, Calibration, Speed and Cost.” 
  2. Release v1.6.1, board revision v1.7.17, benchmarkheaven.com/jev-models, read 7 October 2026. Ranks and scores can change as systems are added or re-measured. 
  3. The fixed reference is Jev 1.13.0 as measured in v1.5: caps of 2× its cost and 2× its median latency. 
  4. Open weights have no public price list, so the board estimates their cost per 1,000 decisions and marks each such figure est. 
  5. The board's published data gives a raw median of 0.029 s, and an adjusted 0.21 s from a fixed “×2 + 0.15 s” allowance it applies to every self-hosted row. Hosted APIs are not adjusted. Latency depends on the GPU. 
  6. The board measured H2O-Lightning-4B v1.1, the text model. Image input is described, with its own measurements, on the model card. 
  7. benchmarkheaven.com/jev-models/api, read 7 October 2026. On the main board, Jev appears as an unranked reference row and other hosted APIs appear only when “Show API offerings” is switched on, also unranked. 
 headshot

Jonathan McKinney

Director of Research

decorator decorator
decorator decorator
h2oai_cube h2oai_cube

Best-in-Class Agents
For Sovereign AI

REQUEST LIVE DEMO