Jul 22, 2026
MOUNTAIN VIEW, Calif., July 22, 2026 — H2O.ai, the world’s leading sovereign enterprise AI platform spanning predictive, generative and agentic AI, today announced that its open-weight H2OVL Mississippi vision-language models have surpassed a combined 2.4 million monthly downloads on Hugging Face. Purpose-built for optical character recognition and document AI, the two models are designed to run within the organization’s own infrastructure, and outperform models many times their size, a shift away from the costly, cloud-bound frontier LLMs that have dominated enterprise AI.
The pull is economic and reflects a broader shift in Enterprise AI. Reading documents with a general-purpose frontier model often means paying per token to send sensitive pages to a frontier LLM provider’s cloud. A purpose-built small model does the same work for a fraction of the cost, on the customer’s own infrastructure, with the data never leaving their walls. What makes the numbers notable is that neither model is new: the 800M model published in December 2024 and the 2B in 2025, and both still clear a million downloads every month. That points to sustained demand, not a launch-driven spike.
"The large model is becoming a commodity. The value is in small, purpose-built models you own and run where your data lives, not rented by the token from someone else’s cloud. Millions of downloads a month tell you where the market is moving. Download it, point it at your own documents, and watch your token bill fall. That’s why we put it in the open."
Cost. Small models turn per-token API bills into fixed, owned infrastructure.
Privacy. Documents stay inside the customer’s environment, which matters most in regulated industries.
Deployability. At under two billion parameters, the models run where the data already lives, whether on-premises, in a private cloud, or air-gapped.
Speed. Smaller, more focused LLM models can have shorter inference response times.
H2OVL Mississippi is a family of two open-weight vision-language models, both released under the Apache 2.0 license and built on H2O.ai’s H2O Danube language models.
H2OVL-Mississippi-2B is a high-performing, general-purpose vision-language model with 2 billion parameters, built for image captioning, visual question answering, and document understanding while staying efficient enough for real-world deployment. Trained on 17 million image-text pairs, it ranks competitively among 2B-class models, leading its size class on MathVista (56.8) and scoring 782 on OCRBench, within a point of the larger Qwen2-VL-2B (797).
H2OVL-Mississippi-800M is a compact 0.8-billion-parameter model trained on 19 million image-text pairs focused on OCR, document comprehension, and chart, figure, and table interpretation. On the text-recognition segment of OCRBench, it scores 274 out of 300, ahead of models many times its size, including the 26-billion-parameter InternVL2-26B.
Text Recognition Results on OCRBench — H2OVL Mississippi-0.8B (274) tops all SLMs and larger LLMs
H2OVL Mississippi-0.8B leads on OCRBench text recognition, beating models up to 30x its size. Source: H2O.ai. Benchmark set as of model release; 274/300 across six text-recognition categories.
Hugging Face is where practitioners discover, test, and benchmark models. H2O.ai is where regulated enterprises bring them into production, with the governance, deployment, and support that mission-critical document workloads require. The efficiency is not theoretical: H2O.ai helped a large North American telecom operator reduce token costs by 10x by moving from a frontier model approach to a small language model fine-tuned on the company’s own data and hosted on its own hardware.
The download numbers reflect a broader shift: as the cost of frontier tokens meets the reality of enterprise-scale document processing, purpose-built small models are becoming the default for the work that does not need a giant model.
H2OVL-Mississippi-2B — OpenVLM benchmark comparison (2B class)
Model | Params | Avg | MMStar | MathVista | OCRBench | MMBench |
|---|---|---|---|---|---|---|
Qwen2-VL-2B | 2.1B | 57.2 | 47.5 | 47.8 | 797 | 72.2 |
H2OVL-Mississippi-2B | 2.1B | 54.4 | 49.6 | 56.8 | 782 | 64.8 |
InternVL2-2B | 2.1B | 53.9 | 49.8 | 46.0 | 781 | 69.6 |
H2OVL-Mississippi-800M — OCRBench text recognition 274/300 · composite OCRBench 751 · trained on 19 million image-text pairs.
Model collection:
https://huggingface.co/collections/h2oai/h2ovl-mississippi-66e492da45da0a1b7ea7cf39
H2O.ai is on a mission to democratize AI for Good. As the world’s leading agentic AI company, H2O.ai converges Generative and Predictive AI to help enterprises and public sector agencies develop purpose-built Agents, SLMs, and solutions on their private data. With a focus on secure, compliant, and infrastructure-flexible Sovereign AI deployments, H2O.ai delivers solutions that align with the highest standards of data privacy and control.
Its open-source technology is trusted by over 20,000 organizations worldwide, including more than half of the Fortune 500. H2O.ai powers AI transformation for companies like AT&T, Commonwealth Bank of Australia, Wells Fargo, Bank of America, Workday, Progressive Insurance, and NIH.
Through its AI for Good program, H2O.ai supports nonprofits, public sector agencies, and global institutions in advancing healthcare, education, disaster response, and environmental sustainability using agentic and predictive AI.
H2O.ai has raised $256 million from investors, including Commonwealth Bank, NVIDIA, Goldman Sachs, Wells Fargo, Capital One, Nexus Ventures, and New York Life.
For more information, visit www.h2o.ai.
Marketing Director • bruna.smith@h2o.ai
Read More