Meta × AWS Bombshell | The Full Story Behind Graviton5's Tens of Millions of Cores
機械翻訳 / Machine-translated

機械翻訳 / Machine-translated
@aifriends
AI Friends(https://aifriends.jp)のクロスポスト公式アカウント。AIツールの紹介・使い方・できることを、中学生でもわかるやさしい日本語で届けます。
"I always thought AI meant GPUs — and then, before I knew it, CPUs had stepped into the spotlight too." On April 24, 2026, the partnership announced by Meta and Amazon Web Services (AWS) is news that symbolizes exactly that kind of turning point. Meta is adopting AWS Graviton5 at a scale of tens of millions of cores, under a multi-year contract worth billions of dollars in total.
Let's unpack the questions — "Why CPUs instead of GPUs?" "What makes Arm so impressive?" "Does any of this matter for Japanese companies?" — in plain language anyone can understand.
First, let's organize the announcement in five minutes.
On April 24, 2026, Amazon and Meta jointly published the partnership across official blogs and press releases.
The two core points: "Meta will deploy AWS Graviton5 processors at a scale of tens of millions of cores," and "the contract spans multiple years and is multibillion-dollar in scale."
"Graviton5 = a server-grade CPU based on Arm, designed in-house by AWS" — currently one of the most powerful cloud-oriented Arm chips on the market.
The fact that "a major tech company is purchasing another company's CPUs in enormous volume under a long-term contract" immediately became a major talking point across the industry.
It's the moment when the behind-the-scenes land grab for CPUs — playing out in the shadow of the GPU arms race — stepped into the open.
Meta's infrastructure chief Santosh Janardhan commented that "diversifying compute sources is now a strategic imperative."
The backdrop is a sense of urgency: "If you depend on a single company's chips, you're at their mercy on pricing negotiations and supply risk alike."
It's the same logic as a logistics strategy of "not relying on just the highway — having bullet trains, planes, and ships ready too."
His statement makes it clear: Meta has made a deliberate corporate decision to absorb the anxiety of the era of Nvidia's dominance through distributed reliance on Arm-based CPUs.
AWS VP and Distinguished Engineer Nafea Bshara said AWS would "provide the foundation for building AI that understands, predicts, and scales efficiently."
This amounts to AWS itself explicitly stating the recognition that "GPUs are the star of training, but running the model — inference — is a different matter entirely."
Think of it as: "The chef who develops new recipes (GPU) and the front-of-house staff who serve large volumes of dishes every day (CPU) require different capabilities."
It is framed as a logical choice for an era in which it is said that 80% of cloud AI operating costs go toward inference.
Here are three angles on why GPU dominance is no longer absolute.
For training large-scale AI models, Nvidia-based GPUs still hold overwhelming strength. But at the stage of actually using a finished AI — inference — situations where CPUs take the lead are rapidly multiplying.
"Learning to cook is training in the kitchen (GPU); actually serving customers every day is running the kitchen operations (CPU)" — a division of labor.
Behind the flashy world of training models with trillions of parameters, the quiet work of inference — repeated hundreds of millions of times — has in reality ballooned to many times the volume of training itself.
AI agents (AIs that autonomously carry out tasks) are characterized by their role as an orchestrator — calling multiple tools and APIs in sequence.
Tasks like "check today's weather, book a restaurant, and calculate the travel route" require a chain of decisions and actions.
"Plating a single dish beautifully is the chef's skill (GPU); managing the flow of an entire multi-course meal is the job of the front-of-house manager (CPU)."
The division of labor — where GPUs excel at matrix computation, while complex branching, state management, and scheduling are where CPUs shine — is becoming the new common sense of the agent era.
Until now, AI infrastructure investment was a competition over "peak performance (FLOPS)." But in an agent era where inference runs around the clock, total cost of ownership (TCO) and power efficiency over 24/7/365 operation become decisive.
It's the same logic as the logistics industry, where "a truck with excellent fuel economy and maintainability" is chosen over the fastest sports car.
The calculation that it's more economical to offload inference to cheap, power-efficient CPUs rather than running expensive GPUs full-throttle on inference — that calculation has clearly held up at Meta's scale, and this contract is the result.
Let's break down the internals of today's strongest class of cloud Arm CPU in plain terms.
Graviton5 packs 192 Arm Neoverse V3 cores — that's 192 high-performance P-cores (performance cores). When you consider that a typical home PC CPU has 8 to 16 cores, this is over 24 times the scale.
"If your home kitchen is a one-person setup, Graviton5 is a hotel kitchen with 192 cooks working simultaneously" — that's the sense of scale.
A single chip can handle hundreds of AI agent sessions in parallel, dramatically raising the number of simultaneous inference connections.
Graviton5 is manufactured on a cutting-edge 3-nanometer (3nm) process, with cache capacity five times that of the previous generation and inter-core communication latency reduced by up to 33%.
"192 cooks, each carrying a huge notepad (cache), passing messages to the cook next to them 30% faster." This means that even in multi-tenant environments where multiple inference jobs run concurrently, performance degradation is minimal — a key advantage.
It is evaluated as approaching the ideal form of a server CPU: balanced, with strong throughput and response time alike.
Performance is up to 25% higher than the previous Graviton4, memory steps up to the ultra-fast DDR5-8800, and I/O supports PCIe Gen6.
"The bandwidth for the data reads and writes that AI agents require expands by an order of magnitude" — that is the design intent.
"Switching from a mountain road (DDR4) to a highway (DDR5-8800)" — data wait times shrink dramatically.
Even if CPU performance improves, it's meaningless if data supply can't keep up. This generation is defined by addressing that bottleneck all at once with state-of-the-art memory and I/O.
Previous Arm-based Graviton chips already had a track record of roughly 20% lower cost and up to 60% less power compared to Intel/AMD x86. Graviton5 is the latest evolution of that lineage — for always-on workloads like AI inference, it holds an overwhelming TCO advantage.
"Handle the same work with half the electricity bill and cheaper chips."
With data center power supply increasingly strained, the era has arrived where CPU selection changes the entire operating cost of a data center.
Here's a look at the in-house Arm CPU strategies of the three major cloud providers.
Google's in-house Arm CPU "Axion," based on Neoverse V2/V3, achieves AMD EPYC Genoa-class thread performance. In benchmarks, some results show it up to 47% faster than Graviton4 on certain metrics.
"A fast middle-distance runner (Axion) vs. a large bus driver (Graviton)" — that kind of relationship.
Speed matters for lightweight inference jobs; core count matters for massive parallel agent workloads — and this defines the segmentation. In a fiercely contested space where the winner can change depending on use case, the current positioning is Axion as performance leader and Graviton as scale leader.
Microsoft's in-house Arm CPU "Cobalt 200" is also based on Neoverse V3, delivering a 50% performance improvement over Cobalt 100. It has been announced that select Azure VM families will begin offering it from early 2026.
The strategy is "strong in single-threaded database workloads, emphasizing TCO optimization across the board."
Three distinct strategies — AWS's scale, Google's speed, Microsoft's TCO — are creating a multipolar Arm CPU market.
Behind the headlines about Nvidia GPU dominance, a fierce technology competition is simultaneously underway on the CPU side.
Intel and AMD in the x86 camp are investing in dedicated AI accelerators (Gaudi, Instinct) and high-efficiency cores (E-cores) to defend their share.
"A long-established local ramen shop reforming its menu in response to a new chain's offensive."
x86 has the advantage in compatibility and leveraging existing assets; Arm has the advantage for new workloads and large-scale deployments — that's the color divide.
The landscape over the next five years will be determined by a critical turning point: which side AI agent workloads end up being optimized for.
Answering the question: "This is happening overseas — what does it have to do with Japan?"
In a joint experiment with NEC, NTT Docomo adopted AWS Graviton2 for its 5G core network infrastructure and succeeded in reducing average power consumption by 72% compared to an x86 environment. In March 2026, it became the first company in Japan to launch commercial operation of a 5G core on AWS.
As a sign that "a critical domestic infrastructure has shifted to Arm-based technology," this is an extremely symbolic case.
With the further efficiency gains of Graviton5, the telecommunications industry's view of it as a tailwind directly addressing power strain in the 5G/6G era is growing.
Optimizing inference costs is the greatest challenge even for teams developing Japan-born LLMs and AI agents.
It is known within the industry that "development teams at UTokyo Matsuo Lab, CyberAgent, PFN, ELYZA, and others are advancing inference benchmarks on Graviton-based systems."
"Combining models that only Japan can produce with infrastructure that can be used cheaply and extensively" is the deciding factor for competitiveness.
The tectonic shift of domestically developed AI moving into the blue ocean of Arm CPU inference — as opposed to the red ocean of competing for Nvidia GPUs — is expected to accelerate in the wake of this announcement.
Migration to Graviton-based instances offered in the AWS Tokyo Region is already underway even at Japanese companies outside the major players.
"20% lower cost, 60% less power" compared to x86 spec differences translate directly into financial statement impact.
"Gradually converting in-house primary web servers to Arm, starting from the top" is the realistic roadmap.
The view is spreading that implementing AI agent features using the Graviton5 generation will become a differentiator for Japanese companies.
Nagata, an SRE at a mid-sized e-commerce site, was agonizing over the monthly AWS invoice.
"Running a small LLM for product recommendations on GPU instances ran over ¥3 million a month." With the partnership as a trigger, he began exploring a phased migration of inference servers to Graviton5-based instances.
"Going from eating high-end sushi every day to home cooking — and finding that satisfaction is almost the same."
The design skill to distinguish where GPUs are necessary from where CPUs are sufficient is emerging as a new competitive arena for SRE professionals.
Kuroki, CTO of a business automation startup, is developing an in-house task automation agent.
"Calling multiple SaaS APIs in sequence, interpreting the results, and deciding the next action" — a textbook agent design. This workload is almost entirely CPU work — the inference core uses a GPU, but the orchestration logic is fine on a CPU.
"With Graviton5, we can handle twice the concurrent connections on the same budget" — that's what the estimate showed.
Arm CPU selection as a differentiator that enables delivering a faster, cheaper agent than competitors has moved to the center of startup growth strategy.
Fukuhara, IT department manager at a major manufacturer, is building an in-house AI assistant.
It's an internal SaaS type that "runs inference continuously for day-to-day inquiry handling." Projecting the inference infrastructure unified on Graviton5 in light of the partnership, the three-year TCO calculation showed a 45% reduction.
"That freed-up budget meant we could increase the dedicated in-house AI education team headcount" — leading to a strategic pivot.
In an era where electricity bills and data center cooling costs are management concerns, the trend of CPU selection becoming an evaluation metric for IT department heads is spreading.
A. No — it's a division of labor, not a replacement. Nvidia-based GPUs hold an overwhelming advantage for training large-scale models; CPUs are more efficient for inference and agent orchestration.
"Assembly of a new car is done by robots; handing it over to the customer is done by a person" — that kind of division.
It's not a substitute for GPUs — it's a partner that takes on the work GPUs aren't suited for. Meta itself has stated clearly that it will use a hybrid strategy combining Nvidia GPUs, its own ASICs, and Graviton5.
A. Most major languages and frameworks already support Arm. Python, Node.js, Go, Java, and Rust have official support; Docker and Kubernetes also run without issue.
"Just a different kind of passport — the countries you can visit are the same." However, legacy apps with x86-only binaries or software with specific driver dependencies will need verification.
Follow the standard procedure: confirm cross-builds in CI/CD before migration and run benchmarks in staging before production — and in most cases the migration will go smoothly.
A. General availability as major EC2 instance families (M8g, C8g, R8g, etc.) is expected within 2026. Rollout to the Tokyo Region will be phased, with major US regions launching first.
"Like a new iPhone arriving in Japan slightly after the US."
The recommended practical approach for Japanese companies: develop and validate in a US region first, then migrate to the Tokyo Region once it becomes available there.
A. A three-step approach is realistic. First, make CI/CD Arm-compatible (set up cross-builds and tests). Second, use AWS Compute Optimizer to visualize how well current workloads are suited for Graviton. Third, design a clear split between the AI inference layer that should stay on GPUs and the parts that should move to Graviton5.
"Sort through your belongings before calling the moving company" — same procedure.
Rather than attempting a full migration all at once, the proven roadmap is to standardize on Graviton starting with new development first.
A. In the short term, there will likely be some wariness as a symbol of the GPU monopoly being cracked. However, Nvidia continues to be the leader in training GPUs and is also competing with inference-dedicated chips (GB200/GB300, Rubin series) — that hasn't changed.
"Even when a new beverage appears in a market dominated by cola, it doesn't bring cola down."
The prevailing industry view is that as the shift from GPU monopoly to a multi-chip era proceeds, Nvidia's share will thin somewhat — but since the overall market is expanding, revenue will continue to grow.
The conventional wisdom of "AI means GPUs" is beginning to be rewritten, starting with the Meta × AWS partnership. Tens of millions of cores in volume, 3nm / 192-core technology, a multi-year contract worth billions of dollars — all of it is evidence that CPUs have stepped into a leading role in the AI agent era.
Breaking away from Nvidia through diversification, the Arm shift accelerating the move away from x86, the economic transition from peak FLOPS to sustained TCO — these three currents are all concentrated in a single contract. For Japanese companies too, it's an opportunity to extend the Graviton power-efficiency revolution — which NTT Docomo pioneered — into the AI agent era.
"The era in which AI infrastructure selection becomes a management decision that directly affects the financial statements" — we are at its threshold right now. Seen that way, the full weight of this announcement comes into focus.
This article is a cross-post from AI Friends.