Ask a European founder to name a homegrown answer to OpenAI and the reply is almost always the same: Mistral AI models. In under three years the Paris startup has gone from a slide deck to a multi billion dollar lab shipping a full stack of open weight and commercial systems. This guide walks through every current Mistral AI model, explains what each one is built for, shows where the real numbers come from, plus traces how the company that makes them turned French machine learning talent into one of the few credible non American frontier labs. If you are choosing an engine for a product or just trying to keep the lineup straight, here is the map.
- Mistral AI ships a layered lineup: open weight flagships like Mistral Large 3, commercial frontier models like Mistral Medium 3.5, edge models under the Ministral 3 name, plus specialists for coding, voice, OCR, along with formal proofs.
- The company was founded in 2023 in Paris and has raised roughly $4 billion, reaching a reported valuation near $14 billion after its September 2025 Series C.
- Open weights under Apache 2.0 remain the strategy that sets Mistral apart from OpenAI and Anthropic, whose frontier models stay closed.
- Annual recurring revenue crossed $400 million by early 2025, up from about $20 million a year earlier, with a stated goal of passing $1 billion.
Who builds the Mistral AI models
Mistral AI was founded in 2023 by three researchers who had spent their careers inside the biggest American labs. Arthur Mensch, the chief executive, came from Google DeepMind. Timothée Lacroix (chief technology officer) and Guillaume Lample (chief science officer) both came from Meta, where they worked on the original LLaMA models. Two advisers from the French health startup Alan, Charles Gorintin plus Jean-Charles Samuelian-Werve, round out the founding circle. The company is headquartered in Paris, and its French identity has become part of the pitch to European enterprises and governments wary of routing sensitive data through United States providers.
That origin story is not just color. The founders left comfortable frontier lab jobs on a specific bet: that the next wave of value would come from models teams could actually own, not rent through someone else's API. Every product decision Mistral has made since flows from that thesis, which is why the open weight flagship exists at all when the easier path would have been to keep everything closed and charge per call.
The funding climb behind the Mistral AI models
The funding history reads like a case study in how fast frontier valuations moved. As TechCrunch reported, Mistral opened with a $113 million seed in June 2023 at a $260 million valuation, then Europe's largest seed round. A €385 million Series A followed in December 2023 at a $2 billion valuation. Microsoft added a small convertible investment in early 2024 alongside an Azure distribution deal, which gave Mistral reach into enterprise customers that a young lab could not have built alone. A June 2024 round of roughly €600 million pushed the valuation to $6 billion, and the €1.7 billion Series C in September 2025 lifted it to about €11.7 billion, near $14 billion. Reporting through 2026 pointed to a further raise around $3.5 billion at a valuation above $23 billion. Backers include Andreessen Horowitz, Lightspeed, Nvidia, Salesforce, ASML, General Catalyst, Bpifrance, plus Microsoft.
Money buys compute, and Mistral has been spending it. The company committed roughly €4 billion to data centers in France and Sweden, launched a Mistral Compute platform in 2026, plus acquired the infrastructure startup Koyeb along with the physics AI firm Emmi. Revenue has followed. Annual recurring revenue sat above $400 million by February 2025, up from about $20 million a year earlier, and Mensch has said the plan is an eventual public listing rather than a sale. For a European lab that many assumed would eventually be swallowed by a United States giant, staying independent this long is itself a result.
The generalist Mistral AI models
The core of the catalog is a set of general purpose text and multimodal models, and this is where the naming rewards a little attention. Mistral publishes its full models overview with version stamps that tell you when each variant shipped, so a model marked v25.12 arrived in December 2025 and one marked v26.04 arrived in April 2026.
Mistral Large 3
Mistral Large 3 (v25.12, released December 2025) is the open weight flagship, described by Mistral as a state of the art, general purpose multimodal and multilingual model. The open weight framing matters. Where OpenAI's GPT line and Anthropic's Claude models keep their weights locked behind an API, Mistral ships downloadable checkpoints for its flagship class. According to launch coverage and third party model guides, Large 3 is a mixture of experts design in the range of 41 billion active parameters out of roughly 675 billion total, which would make it one of the largest open weight models released by a major lab. Treat the exact parameter figures as vendor and community reported rather than independently audited, but the shape is clear: a big MoE model you can host yourself. The mixture of experts design is what makes that hostable, since only a fraction of the total parameters activate on any given token, cutting the compute needed to serve it.
Mistral Medium 3.5
Mistral Medium 3.5 (v26.04) is the newer commercial frontier model, and it is not open weight. Mistral positions it as a cost efficient, enterprise focused system tuned for reasoning, coding, plus agentic work, with simplified deployment. This is the model Mistral points enterprise buyers toward when they want frontier quality behind an API without self hosting. The split between an open weight flagship and a closed, cheaper commercial tier is a deliberate business design: give the community the big model to build trust, then monetize the tuned, hosted variant that most enterprises will actually deploy.
Mistral Small 4 and Les Ministraux
Mistral Small 4 (v26.03) is the workhorse of the lineup and ships under an Apache 2.0 license, which permits commercial use without royalty. Mistral calls it a hybrid model that unifies instruction following, reasoning, plus coding in one system, with multimodal and multilingual support. Community reporting puts it near 119 billion total parameters with only a few billion active per token thanks to the mixture of experts routing, and Mistral has listed input pricing as low as $0.15 per million tokens for the hosted version. That price point is the real story: a permissively licensed model cheap enough to run at scale, which is exactly the profile most production teams want for high volume work.
Below Small sit the edge models. Ministral 3 comes in 14B, 8B, plus 3B sizes (all v25.12), branded as Les Ministraux and aimed at on device and latency sensitive workloads where a giant model is overkill. These are the models you reach for when the target is a phone or a private server rather than a data center, and their small footprint means they can run without a cloud connection at all.
Coding and voice: Mistral AI models built for one job
Beyond the generalists, Mistral runs a bench of specialists, and several are genuinely best in class for their niche.
On code, Devstral 2 (v25.12) is the open weights agentic coding model, built for autonomous software engineering rather than single line completion. Earlier Devstral releases were reported at 123 billion parameters with a 256,000 token context window and SWE-bench scores in the low seventies, and the v25.12 refresh continues that agentic focus. Sitting alongside it is Codestral (v25.08), a low latency completion engine tuned for fill in the middle and high frequency code generation across more than 80 programming languages. The division of labor is clean. Codestral handles the fast in editor suggestions that need to return in milliseconds, while Devstral drives the agent that opens files, edits them, then runs tests to check its own work.
The audio stack is newer. Voxtral TTS (v26.03) is Mistral's text to speech model with zero shot voice cloning and multilingual support, a direct shot at ElevenLabs and OpenAI's voice offerings. Voxtral Mini Transcribe Realtime (v26.02) handles low latency transcription, and Voxtral Small covers heavier speech to text work. For documents, OCR 4 extracts text with paragraph level bounding boxes, a capability that matters for anyone feeding scanned contracts or forms into a pipeline where layout, not just raw text, carries meaning.
The most unusual specialist is Leanstral 1.5, an open source agent built for Lean 4 formal proof engineering. It generates both code plus a machine checkable mathematical proof that the code is correct, a narrow but serious tool for formal verification work. Rounding out the roster are embedding and safety models such as Codestral Embed, Mistral Embed, plus Mistral Moderation 2. The current lineup is visible on Mistral's models page, which groups everything from cloud to edge.
Version stamps like v25.12 and v26.04 encode the release year and month. When you see two models with similar names, the stamp tells you which one is newer, which is the fastest way to avoid pinning an outdated variant in production.
Why open weights are the Mistral AI models' real edge
Every lab has a coding model and a voice model now. What still separates Mistral is the willingness to release weights. Apache 2.0 licensing on models like Small 4 lets a company download the checkpoint, run it on its own hardware, fine tune it on private data, then ship it in a product without asking permission or paying a per token toll. For regulated industries, defense adjacent work, plus any team that cannot send data to a United States API, that is not a nice to have. It is the entire reason to pick Mistral.
Mensch has leaned into this argument publicly. In commentary picked up by The Decoder, he argued that proprietary models give the labs behind them a front row seat to a customer's business processes, and that open weights are the way to keep that visibility out of a competitor's hands. Whether or not you buy the framing, it is a coherent position the closed labs cannot match, because their business depends on the weights staying private. If you are weighing open models for a commerce or content product, the same logic shows up in how an open-source AI stack keeps prompts off the big proprietary providers.
The strategy is not free of tension. Open weights make it harder to charge for the flagship, which is why Medium 3.5 exists as a closed, hosted tier. The company is threading a needle: release enough to stay the open lab of record in Europe, while keeping a commercial model good enough that enterprises pay for the managed version. So far the revenue trajectory suggests the needle is holding, and the willingness of large customers to sign multi year deals suggests they see the open weight guarantee as a feature rather than a risk.
How the Mistral AI models compare to the American labs
On raw benchmark leaderboards, Mistral's flagships trail the very top American frontier models on the hardest reasoning and coding tests. That gap is real and worth stating plainly. What Mistral offers instead is a different trade: near frontier quality you can actually download, at prices well under the closed leaders, from a company outside United States jurisdiction. For a large slice of production workloads, especially retrieval, classification, extraction, plus mid difficulty generation, that trade is a win even when a closed model would score a few points higher on paper. The right question is rarely which model tops the leaderboard. It is which model clears your quality bar at a price and a governance profile you can live with.
The assistant side tells a similar story. Le Chat, Mistral's consumer facing chat product, hit one million downloads in fourteen days after a push, yet it still trails ChatGPT badly on brand recognition. Mistral's advantage was never going to be the consumer app. It is the model layer underneath, sold to developers and enterprises who care about ownership, price, and jurisdiction more than about a household name.
The partnership list backs that up. Mistral has signed distribution or deployment deals with Microsoft, Nvidia, ASML, IBM, Orange, Stellantis, the shipping group CMA CGM, plus the news agency Agence France-Presse, alongside French and European government work. These are infrastructure and enterprise relationships, not viral growth. They are also exactly the customers who benefit most from open weights plus European hosting, which is why Mistral keeps investing in the model layer rather than chasing consumers.
Choosing among the Mistral AI models
If you are mapping the lineup to a real decision, a few rules of thumb hold. For a self hosted general model with a permissive license, start with Mistral Small 4 and move up to Mistral Large 3 only if you need the extra headroom and can afford the hardware. For a hosted enterprise deployment where you want frontier quality without running the infrastructure, Mistral Medium 3.5 is the target. For anything on device or latency bound, the Ministral 3 sizes cover phones through private servers.
On code, pair Codestral for fast in editor completion with Devstral 2 when you need an agent that can navigate a repository on its own. For voice, Voxtral TTS plus the Voxtral transcription models now cover both directions of the audio pipeline, while OCR 4 handles document extraction. Leanstral is a specialist you will know you need only if you are doing formal verification, and most teams never will. The point of listing it here is to show the breadth of the bench. Mistral is not just shipping a chat model with a coding variant bolted on. It is building a portfolio wide enough that a single vendor relationship can cover text, vision, speech, documents, plus verification, which is a meaningful pitch for an enterprise that would rather not stitch together five providers. The full set of version stamps and capability tags lives in Mistral's documentation plus its ongoing news feed, which is where new releases land first.
A European bet worth watching
The through line across every Mistral AI model is a wager that open weights, competitive pricing, plus European sovereignty add up to a durable business even against labs with more capital and higher benchmark peaks. The 2026 spending, the data center buildout, along with the revenue climb toward a billion dollars, suggest the wager is paying enough to keep the company independent and, by Mensch's own account, aimed at a public listing rather than an acquisition. For anyone building on top of language models, the practical takeaway is simpler. Mistral gives you a real second source, one you can download and host, and that optionality is valuable regardless of which model wins the next benchmark. Keep an eye on the promised open weight release later in 2026, because that is where the next chapter of this lineup gets written, and it will show whether the open strategy can keep pace as the frontier keeps moving.