GSMA Open-Telco LLM Benchmarks 2.0

The telecom LLM benchmark.
And where TSLAM stands.

The GSMA tested 23 leading AI models on real telecom tasks. TSLAM-G3 ranked first for domain reasoning, proving focused models can outperform far larger general systems.

1stDomain reasoning (TeleQnA), score 82.5 — TSLAM-G3
3rdTeleMath & 3GPP-TSG — TSLAM-18B
23Models benchmarked, including the largest frontier labs
30×Smaller than the 600B+ generalists it competes with

The benchmark

What the GSMA actually measures

The Open-Telco LLM Benchmarks test telecom knowledge, network maths and 3GPP standards, giving operators a neutral measure based on the work their teams perform.

Telecom domain knowledge & reasoning

TeleQnA

Thousands of questions spanning standards, architectures, protocols and regulatory complexity — the core test of whether a model actually understands telecom.

Quantitative network engineering

TeleMath

Mathematical reasoning on real network problems: capacity, propagation, traffic and performance calculations.

Standards comprehension

3GPP-TSG

Mastery of 3GPP technical specification group documents — the 4G/5G standards that define how networks are built.

Results · November 2025

Three TSLAM models, three validated strengths

Independent results show where each specialized model leads: domain reasoning, standards-aware engineering and efficient edge deployment.

Telecom knowledge copilot

1st

Domain reasoning (TeleQnA) — score 82.5

TSLAM-G3

Topped the GSMA leaderboard in Domain Reasoning (TeleQnA) with a score of 82.5, ahead of the largest frontier-lab models — validated as the industry's premier solution for telecom domain intelligence.

GSMA independently validated

Telecom engineering brain

3rd

TeleMath & 3GPP-TSG

TSLAM-18B

Ranked #3 in TeleMath and 3GPP-TSG comprehension, within 5% of the top-ranked 600B+ parameter generalist model — while being roughly 30× smaller, and deployable on-premise.

GSMA independently validated

Telecom AI at the edge

Edge

Cited by the GSMA for energy efficiency

TSLAM-2B Mini

The GSMA report highlighted it as delivering “respectable accuracy, while maintaining far lower energy requirements for sustainable deployment” — precision-built for on-device, low-latency tasks.

GSMA independently validated

Why it matters

Specialization beats scale on telecom tasks

A telecom model about 30 times smaller than leading general systems can deliver strong accuracy with lower cost, local control and sovereign deployment.

  • Deployment freedom — edge, on-premise and hybrid, not cloud-only.
  • Data sovereignty — sensitive network and customer data never leaves your perimeter.
  • Lower TCO & energy — compact models cut inference cost and carbon footprint at operator scale.

FAQ

Telecom LLM benchmarks, answered

Clear answers on the methodology, rankings, specialization and independent verification behind the results.

What is the GSMA Open-Telco LLM Benchmark?

An independent, industry-neutral benchmark run by the GSMA — the global association of mobile operators — that evaluates large language models on telecom-specific tasks such as domain reasoning (TeleQnA), quantitative network engineering (TeleMath) and 3GPP standards comprehension (3GPP-TSG). Version 2.0 evaluated a field of 23 models, including the largest general-purpose frontier models.

How did NetoAI's TSLAM models rank?

TSLAM-G3 topped the Domain Reasoning (TeleQnA) leaderboard with a score of 82.5. TSLAM-18B ranked #3 in TeleMath and 3GPP-TSG, within 5% of the top 600B+ parameter generalist model while being roughly 30× smaller. TSLAM-2B Mini was cited by the GSMA for delivering respectable accuracy at far lower energy requirements, making it suitable for edge deployment.

Why do specialized telecom LLMs beat much larger general-purpose models?

Networks are physics — RF propagation, photonics, electromagnetism — and telecom workflows are encoded in standards and operational data that barely exist on the public web. TSLAM is trained on 100B+ tokens of engineered telecom-specific data, so it reasons about networks the way an engineer does, at a fraction of the size, cost and energy of frontier generalist models.

Can I verify or reproduce these results?

Yes. The GSMA publishes the benchmark and leaderboard independently, and NetoAI open-sources TSLAM variants and the NetBench dataset on Hugging Face, so operators and researchers can run their own evaluations.

Built for your network

Benchmark TSLAM on your own network data

Schedule a demo
Schedule a demo

Tell us a little about your network and our team will tailor the walkthrough.