Telecom domain knowledge & reasoning
TeleQnA
Thousands of questions spanning standards, architectures, protocols and regulatory complexity — the core test of whether a model actually understands telecom.
The GSMA tested 23 leading AI models on real telecom tasks. TSLAM-G3 ranked first for domain reasoning, proving focused models can outperform far larger general systems.
The benchmark
The Open-Telco LLM Benchmarks test telecom knowledge, network maths and 3GPP standards, giving operators a neutral measure based on the work their teams perform.
Telecom domain knowledge & reasoning
Thousands of questions spanning standards, architectures, protocols and regulatory complexity — the core test of whether a model actually understands telecom.
Quantitative network engineering
Mathematical reasoning on real network problems: capacity, propagation, traffic and performance calculations.
Standards comprehension
Mastery of 3GPP technical specification group documents — the 4G/5G standards that define how networks are built.
Results · November 2025
Independent results show where each specialized model leads: domain reasoning, standards-aware engineering and efficient edge deployment.
Telecom knowledge copilot
Domain reasoning (TeleQnA) — score 82.5
Topped the GSMA leaderboard in Domain Reasoning (TeleQnA) with a score of 82.5, ahead of the largest frontier-lab models — validated as the industry's premier solution for telecom domain intelligence.
Telecom engineering brain
TeleMath & 3GPP-TSG
Ranked #3 in TeleMath and 3GPP-TSG comprehension, within 5% of the top-ranked 600B+ parameter generalist model — while being roughly 30× smaller, and deployable on-premise.
Telecom AI at the edge
Cited by the GSMA for energy efficiency
The GSMA report highlighted it as delivering “respectable accuracy, while maintaining far lower energy requirements for sustainable deployment” — precision-built for on-device, low-latency tasks.
Why it matters
A telecom model about 30 times smaller than leading general systems can deliver strong accuracy with lower cost, local control and sovereign deployment.
Sources
Review the official leaderboard, independent reporting, peer-reviewed research and open models behind every result shown here.
Leaderboard · GSMA
Open →
Press · TechStory
Open →
White paper · GSMA
Open →
Models · Hugging Face
Open →
Research · EMNLP 2025
Open →
Initiative · GSMA × NetoAI
Open →
FAQ
Clear answers on the methodology, rankings, specialization and independent verification behind the results.
An independent, industry-neutral benchmark run by the GSMA — the global association of mobile operators — that evaluates large language models on telecom-specific tasks such as domain reasoning (TeleQnA), quantitative network engineering (TeleMath) and 3GPP standards comprehension (3GPP-TSG). Version 2.0 evaluated a field of 23 models, including the largest general-purpose frontier models.
TSLAM-G3 topped the Domain Reasoning (TeleQnA) leaderboard with a score of 82.5. TSLAM-18B ranked #3 in TeleMath and 3GPP-TSG, within 5% of the top 600B+ parameter generalist model while being roughly 30× smaller. TSLAM-2B Mini was cited by the GSMA for delivering respectable accuracy at far lower energy requirements, making it suitable for edge deployment.
Networks are physics — RF propagation, photonics, electromagnetism — and telecom workflows are encoded in standards and operational data that barely exist on the public web. TSLAM is trained on 100B+ tokens of engineered telecom-specific data, so it reasons about networks the way an engineer does, at a fraction of the size, cost and energy of frontier generalist models.
Yes. The GSMA publishes the benchmark and leaderboard independently, and NetoAI open-sources TSLAM variants and the NetBench dataset on Hugging Face, so operators and researchers can run their own evaluations.
Built for your network