AT&T and Microsoft scale trillion-token workloads with Microsoft Foundry and AMD


Telecommunications organizations are more and more trying to AI to assist groups navigate extremely specialised domains, however generic fashions usually lack the industry-specific data wanted to grasp telecom networks, requirements, and operations. To deal with that hole, AT&T created their Open Telco (OTel) fashions, the following era of telecom-focused AI designed to convey deeper telecommunications experience into AI methods. Constructing OTel2.0 required greater than coaching a big language mannequin, it mirrored a broader challenge many organizations face: how you can construct domain-specific AI methods at scale whereas balancing value, efficiency, and operational complexity. Price administration rapidly grew to become a key consideration. To proceed advancing telecom-focused AI, AT&T wanted a platform able to supporting OTel2.0 improvement at a completely new scale.

The place groups beforehand needed to personal and handle deployments, infrastructure, and the related operational overhead, Foundry Managed Compute supplied a extra streamlined solution to entry devoted graphics processing unit (GPU) capability. This transformation requires greater than highly effective fashions; it requires the power to scale with out compromising value, flexibility, or efficiency.

Utilizing Microsoft Foundry Managed Compute, AT&T was capable of experiment throughout a number of open fashions, optimize workloads throughout totally different GPU architectures, and course of large volumes of telecom knowledge all inside a unified platform. The outcome was an AI improvement setting able to supporting trillions of tokens whereas giving groups the flexibleness to iterate, optimize, and innovate sooner.

Mannequin selection meets infrastructure flexibility

Constructing OTel2.0 required flexibility throughout each fashions and infrastructure. Slightly than standardizing on a single mannequin, AT&T adopted a multi open-model technique. Open fashions had been central to AT&T’s method as a result of they supplied the flexibleness to work with authorized telecom knowledge, tailor the workflow for domain-specific mannequin improvement, and assist large-scale experimentation with better management over value and deployment technique. Via Microsoft Foundry, the workforce deployed a number of fashions from the Hugging Face assortment, together with Phi-4, OSS-120B, and Gemma-4, to assist totally different levels of improvement, from artificial knowledge era and knowledge preparation to reasoning-intensive workloads and broader mannequin improvement efforts. Phi-4 performed a big position on this course of, processing greater than 700 billion tokens a month as a part of the broader knowledge preparation and coaching workflow for OTel2.0.

Each firm on the planet must construct its personal AI, and that’s solely attainable with open fashions and open supply. AT&T is championing this imaginative and prescient, constructing on open fashions like Phi-4 and Gemma, and giving OTel again to the neighborhood as a telecom AI basis others can construct upon. Microsoft Foundry makes this sensible at scale, bringing the newest open fashions from the Hugging Face assortment along with AMD and NVIDIA GPUs in a single place, so groups can choose the correct mannequin and the correct {hardware}, then deploy in hours as a substitute of weeks.

—Jeff Boudier, Vice President of Product, Hugging Face

Creating OTel2.0 additionally required infrastructure able to working at telecom scale. AT&T used roughly 530 GPUs by way of Microsoft Foundry Managed Compute spanning a number of GPU architectures together with 430 AMD Intuition™ MI300X GPUs. This heterogenous method gave AT&T extra flexibility in how fashions had been deployed and optimized as necessities advanced.

Desk 1: Explains what open supply fashions had been used and the way

This flexibility illustrates a broader development throughout AI improvement. Organizations more and more want platforms that enable them to decide on the correct mannequin for the job, optimize for value and efficiency, and scale workloads with out rebuilding operational environments. Microsoft Foundry brings mannequin selection, infrastructure flexibility, governance, and operational scale collectively in a unified platform that helps these necessities.

Past flexibility and price, deployment velocity is a essential issue for a lot of AI initiatives. As workloads increase and new fashions are evaluated, the power to entry GPU capability rapidly allows groups to maneuver from experimentation to execution sooner with out prolonged provisioning cycles. With Foundry Managed Compute, AT&T may deploy and scale fashions in days quite than ready weeks for infrastructure to grow to be obtainable, serving to speed up improvement timelines and preserve momentum throughout OTel2.0 improvement.

Optimizing value with out limiting innovation

As AI workloads develop, economics grow to be as vital as mannequin efficiency. For AT&T, one of many major aims was to decrease AI mannequin consumption prices whereas persevering with to drive significant enterprise worth by way of AI-powered innovation. By utilizing open fashions on Microsoft Foundry Managed Compute, AT&T was capable of assist large-scale knowledge preparation and mannequin improvement utilizing a special financial mannequin constructed round devoted GPU infrastructure and open-model flexibility.

The influence grew to become clear at scale. In assist of OTel2.0, AT&T processed roughly 1T tokens, consisting of uncooked paperwork from GSMA supplemented by artificial knowledge generated. Producing the info utilizing open-source fashions like Phi-4, served by Microsoft’s Foundry Managed Compute, saved tens of tens of millions of {dollars} versus utilizing frontier fashions. This allowed groups to spend money on larger-scale experimentation and improvement whereas sustaining a give attention to enterprise worth and operational effectivity.

Desk 2: Fast info concerning the OTel mannequin household and metrics round what was used to construct OTel2.0 

When you’re processing lots of of billions of tokens, infrastructure turns into a part of the issue you resolve. Foundry Managed Compute gave us entry to GPU capability at scale so our groups may give attention to advancing OTel2.0 as a substitute of managing infrastructure.

—Mark Austin, Vice President, Information Science and AI at AT&T

At this scale, infrastructure is not merely a deployment consideration. It turns into a strategic part of AI improvement.

Accelerating the following wave of production-scale AI

OTel 2.0 demonstrates how organizations can mix open fashions, scalable infrastructure, and area experience to construct production-ready AI methods. By matching totally different fashions to totally different workloads and optimizing infrastructure for value and efficiency, AT&T was capable of course of trillions of tokens whereas sustaining operational effectivity. 

As organizations transfer from AI experimentation to manufacturing deployment, they more and more want the flexibleness to decide on the correct fashions, optimize infrastructure, and scale effectively. Microsoft Foundry and Foundry Managed Compute assist assist that transition by bringing these capabilities collectively in a unified platform.

Study extra

Discover session subjects from AMD’s Advancing AI:



Deixe um comentário

O seu endereço de e-mail não será publicado. Campos obrigatórios são marcados com *