Research

Qwen3.7 Max Tops Multilingual Benchmarks as Europe’s Sovereign Models Rise

Alibaba’s Qwen3.7 Max scores 100% weighted multilingual on the September 2026 BenchLM leaderboard, ahead of Claude Opus 4.5, while European and open-weight options reshape enterprise choices.

Editorial·2 Sep 2026
Qwen3.7 Max Tops Multilingual Benchmarks as Europe’s Sovereign Models Rise

Alibaba’s Qwen3.7 Max has taken the top spot for multilingual performance in the September 2026 BenchLM leaderboard, achieving a weighted multilingual score of 100%. Anthropic’s Claude Opus 4.5 ranks second at 82.9%, followed by Qwen3.7 Plus at 78.9%. The rankings, updated on September 2, 2026, are based on two benchmarks: MMLU-ProX, which tests broad professional knowledge across languages, and MGSM, which evaluates multilingual mathematical reasoning.

For global enterprises and public-sector bodies, these figures are not merely technical trivia. The weighted multilingual score contributes 7% to a model’s overall evaluation on BenchLM, but language capability increasingly determines whether an AI system can serve customers, regulators, and employees across borders. A model that excels in English but stumbles in Arabic, Japanese, or Polish can create compliance risks, uneven user experiences, and hidden costs. The September 2026 results also highlight a broader strategic shift: the best multilingual models are no longer concentrated in a single region, and organizations now face a three-way decision among global frontier models, sovereign European providers, and open-weight alternatives.

Qwen’s multilingual lead and the benchmark method

Qwen3.7 Max’s 100% weighted score reflects the highest weighted composite across the two underlying benchmarks. MMLU-ProX tests professional knowledge in fields such as law, medicine, and engineering across multiple languages, while MGSM evaluates mathematical reasoning in a multilingual setting. BenchLM combines these into a single weighted multilingual score. The gap between first and second place is notable: Claude Opus 4.5, despite strong overall performance, trails by 17.1 percentage points on this specific measure. Qwen3.7 Plus, a sibling model from Alibaba, holds third place at 78.9%.

The dominance of Qwen in multilingual tasks underscores the global nature of AI development. Alibaba’s Qwen family has consistently pushed into languages that are often underserved by Western labs, and the September 2026 leaderboard suggests that investment is paying off. For international professionals, the result means that choosing a model for a multilingual product or service cannot rely on brand familiarity alone. A model that leads in English-language benchmarks may lag significantly when evaluated across Arabic, Hindi, Swahili, or Vietnamese.

European providers and the data sovereignty question

For organizations that cannot or will not send sensitive data to US-based cloud providers, European LLM providers have become a critical alternative. A June 2026 analysis by Eden AI identified five leading European providers:

  • Mistral AI (France) — recommended for its balance of quality, open-weight options such as Mistral Large 2 and Mixtral 8x22B, and enterprise flexibility.
  • Aleph Alpha (Germany) — top choice for sovereign deployments in government and regulated industries, with models like Pharia running on EU infrastructure.
  • AMD Silo AI (Finland) — specializes in Nordic and low-resource European languages with fully open-source Poro and Viking model families.
  • LightOn (France) and DeepL (Germany) — additional European providers with distinct enterprise offerings.

These providers matter because data governance is now inseparable from model selection. The EU AI Act imposes requirements on high-risk AI systems, and the US CLOUD Act can compel US-based providers to hand over data even when it is stored abroad. European models hosted on EU infrastructure offer a clearer path to compliance for public-sector bodies, healthcare organizations, and financial institutions. The trade-off is often scale and frontier performance, but for many use cases, regulatory alignment outweighs a few points on a leaderboard.

Open-weight models and the self-hosting calculus

The open-source LLM landscape in 2026 is robust enough to compete with proprietary systems in many enterprise scenarios. Three permissively licensed models stand out:

  • DeepSeek-V4-Pro — 1.6 trillion parameters, MIT license.
  • Qwen3.5-397B-A17B — Apache 2.0 license.
  • GLM-5.2 — 754 billion parameters, MIT license.

Open-weight models allow enterprises to self-host, ensuring data privacy, avoiding vendor lock-in, and enabling fine-tuning for specific domains such as legal contracts, medical records, or technical documentation. This option is particularly valuable for European businesses seeking compliance with the EU AI Act and protection from the US CLOUD Act. Self-hosting removes the question of where data resides and who can access it. However, open models require significant infrastructure investment. Running a 1.6-trillion-parameter model demands substantial GPU capacity, engineering expertise, and ongoing maintenance. Open models often become cost-effective at scale and benefit from rapid community-driven innovation, but the upfront cost and operational burden are real. For smaller organizations, managed open-weight offerings or European API providers may be a more practical middle ground.

What the September 2026 landscape means for decision-makers

The choice of an LLM is no longer just a technical decision, but a strategic one involving data governance, cost structure, and long-term control.

#large language models #multilingual AI #benchmarks #data sovereignty

Newsletter

Get the AI news that matters

One short brief with the day's most important AI stories — written for professionals.

We send a confirmation link. No spam. Unsubscribe anytime.

WhatsApp