DevOps / SRE / Platform · 25.08.2026, 19:31 UTC
IBM’s new Granite 4.2 models add reasoning and stay dense
| Schweregrad | info |
|---|---|
| Kategorie | DevOps / SRE / Platform |
| Quelle | The New Stack ↗ |
| Veröffentlicht | 25.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
On Tuesday, IBM launched the latest family of its open-weight Granite large language models (LLMs). Weighing in at 3 billion, 8 billion, and 30 billion parameters, IBM is taking a very different approach to model building from some of the competition here, with dense, decoder-only reasoning models that it pre-trained from scratch.
Many recent models have moved from all-attention Transformers toward hybrid Mamba/attention architectures, including Nvidia’s Nemotron 3 family. But IBM tried that with its Granite 4.0 models. That generation included conventional dense, dense-hybrid, and hybrid MoE models.
With Granite 4.1, IBM returned the main family to an all-attention, dense Transformer architecture. At the time, IBM said that these new models outperformed the older generation, “while using a simpler — and therefore more flexible — architecture for fine-tuning for downstream tasks.”
Reasoning with Granite
IBM describes the 4.2 family as a “reasoning-focused release.” Models can run in thinking and non-thinking modes, but there is also a low-effort mode that only spends a low number of reasoning tokens for answering easy questions.
When the company launched the Granite 4.1 models, IBM still argued that reasoning models weren’t efficient enough, so “turning to less expensive, non-reasoning models with similar benchmark performance for select tasks like instruction following and tool calling makes sense for enterprise users.” At this point, though, the team clearly believes that reasoning is a necessary feature — even as it remains optional for the 4.2 models.
Unlike …
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info AWS Secrets Manager adds managed external secrets support for Cisco Security Platform and Netskope
- info Towards Understanding On-Policy Distillation through the Lens of Test-Time Scaling
- info Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA
- info Benchmarking and Boosting Multilingual Capabilities of LVLMs via OCR-Centric Reinforcement Learning