In an era where Artificial Intelligence is shifting from simple chat interfaces to autonomous, action-oriented workflows, JetBrains has taken a significant step forward. The company today announced the release of Mellum2.1, an evolution of their 12B mixture-of-experts (MoE) model. Designed specifically to excel in the nuanced, high-stakes environment of software development, Mellum2.1 represents a refined marriage between compact architectural efficiency and advanced, environment-aware reinforcement learning.
For developers and enterprises looking to run AI agents locally—without the latency or privacy concerns associated with massive, cloud-bound models—Mellum2.1 offers a compelling, open-source solution that challenges the dominance of larger, general-purpose LLMs.
Main Facts: The Evolution of a Coding Powerhouse
Mellum2.1 retains the architectural footprint of its predecessor, the original Mellum2 released in June 2026. It remains a 12B parameter mixture-of-experts model, utilizing 2.5B active parameters per inference pass. By maintaining this compact footprint, JetBrains has ensured that the model remains highly performant on consumer-grade and enterprise-edge hardware.
The primary advancement in version 2.1 lies not in the underlying architecture, but in the "post-training" phase. JetBrains engineers spent the summer subjecting the model to rigorous reinforcement learning (RL) in real-world software development environments. By simulating millions of sandboxed runs across thousands of distinct codebases, the team moved the model beyond simple text completion and into the realm of true agentic behavior.
Mellum2.1 is now capable of:
- Navigating Complex Repositories: Understanding file structures and project-wide dependencies.
- Autonomous File Editing: Proactively suggesting and implementing code changes.
- Self-Validation: Iteratively testing and checking its own code output to ensure functional correctness before suggesting a commit.
Licensed under the Apache 2.0 license, this release reinforces JetBrains’ commitment to the open-source ecosystem, inviting developers to integrate the model into their own infrastructure without restrictive proprietary barriers.

Chronology: From Concept to Agentic Capability
The journey to Mellum2.1 is rooted in the broader JetBrains strategy to integrate AI directly into the developer experience.
- June 2026: JetBrains open-sourced the original Mellum2. While highly praised for its speed and efficiency, the model primarily functioned as a sophisticated autocomplete and suggestion engine. Feedback from the developer community highlighted a need for more "agentic" capabilities—the ability to act as a partner in the development process rather than just a tool.
- Summer 2026: The JetBrains AI research team pivoted to an intensive RL-focused training cycle. Utilizing thousands of sandboxed environments, the team allowed the model to "live" within codebases, failing and succeeding in tasks to refine its decision-making logic.
- October 2026: The official release of Mellum2.1. This version encapsulates the lessons learned from those millions of simulated interactions, resulting in a model that understands the recursive, logical nature of programming.
Supporting Data: Performance and Throughput
To validate the efficacy of the new training regimen, JetBrains conducted a series of comparative benchmarks against industry peers, including Qwen3.5-9B and Gemma 4 E4B.
Agentic Proficiency
In agentic coding tasks—which require multi-step reasoning, tool usage, and environmental interaction—Mellum2.1 demonstrated a marked improvement over Mellum2. The model exhibited higher accuracy in complex problem-solving scenarios, competitive programming, and mathematical reasoning. Perhaps most importantly, it showed a consistent performance profile across both simple, everyday coding tasks and high-complexity architectural refactoring.
Latency and Speed Metrics
JetBrains has leaned heavily into the integration of Multi-Token Prediction (MTP) to ensure that the model remains responsive under heavy load. The technical specifications highlight a significant advantage in throughput:
- Throughput: Under stress tests, Mellum2.1 operates as the fastest model in its class, serving nearly twice the number of tokens per second compared to Qwen3.5-9B.
- Speculative Decoding: For single-request latency, the inclusion of MTP makes the model approximately 1.6 times faster than its predecessor.
These metrics are crucial for developers building sub-agents that need to provide real-time feedback within an IDE. A sluggish AI model often breaks the "flow state" of a developer; by prioritizing tokens-per-second without sacrificing reasoning capabilities, Mellum2.1 ensures the AI remains an invisible, high-speed participant in the coding process.
Official Responses and Strategic Vision
Bulat Salimzianov, representing the engineering team behind the model, emphasizes that the goal was never to create a "jack-of-all-trades" model, but rather a "master-of-one."

"Open source is how better models get made," the team noted in the official release documentation. By releasing the model under Apache 2.0, JetBrains is actively soliciting community feedback. The company has framed Mellum2.1 as a foundational component for the next generation of AI tools.
"If you’re building coding agents, sub-agents, or AI tools that run on your own infrastructure, we’d love for you to try Mellum2.1," the statement reads. "Your feedback will shape the next version." This collaborative approach is intended to decentralize the development of coding AI, moving power away from centralized cloud APIs and back into the hands of individual developers and local infrastructure teams.
Implications: The Future of Localized AI Development
The release of Mellum2.1 carries significant implications for the software industry at large.
1. The Rise of "Small" AI
For years, the industry narrative has been that "bigger is better." However, Mellum2.1 proves that by focusing on domain-specific training and efficient MoE architecture, models with a smaller parameter count can outperform their larger, more generalist counterparts in specific tasks. This is a win for sustainability and hardware accessibility.
2. Privacy and On-Premise Security
Many enterprises are restricted by strict data privacy compliance, preventing them from sending proprietary codebase snippets to cloud-based LLMs. By providing a high-performance model that can be run locally (or within a secure, private cloud environment), JetBrains is enabling a broader range of companies to leverage AI agents without risking intellectual property leakage.
3. Democratizing the "Coding Agent"
Previously, building a fully autonomous coding agent required access to expensive, proprietary APIs. With Mellum2.1 being open source and optimized for llama.cpp, Ollama, and LM Studio, the barrier to entry has dropped significantly. Small startups, hobbyists, and research labs can now experiment with agentic workflows that were, until recently, reserved for the largest tech conglomerates.

4. The Iterative Future
The roadmap for Mellum2.1 is clearly focused on community-driven growth. With GGUF builds and MTP heads for vLLM currently in the pipeline, the ecosystem around the model is set to expand rapidly. The focus on "post-training" rather than "pre-training" suggests that JetBrains intends to continue refining the model’s behavior, making it more intuitive and less prone to the hallucinations that plague larger, less focused models.
Conclusion
Mellum2.1 is not just a software update; it is a declaration of intent. By choosing to iterate on a compact, fast, and open-source architecture, JetBrains is championing a future where AI is a deeply integrated, highly responsive, and private utility for every developer.
As the industry watches to see how Mellum2.1 performs in the wild, one thing is clear: the era of the "AI-powered developer" is no longer a distant promise. It is here, it is fast, and thanks to the open-source community, it is getting better every day. Developers interested in testing the model can find it on Hugging Face, where the community is already beginning to build the next wave of agentic tools. Whether you are building a custom sub-agent for a specialized framework or simply looking to enhance your IDE experience, Mellum2.1 stands as a potent, accessible tool for the modern engineer.








