metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
iPhone Duo
Apple A20 Pro
On-Device AI
Apple Intelligence
Neural Engine

iPhone Duo: Apple's A20 Pro and Its On-Device AI Bet (2026)

Apple's iPhone Duo foldable runs on an A20 Pro chip with a dual 16-core Neural Engine. Here is what its on-device AI actually does, and where Apple still relies on the cloud.

Metir AI TeamSeptember 9, 20268 min read
iPhone Duo: Apple's A20 Pro and Its On-Device AI Bet (2026)

On September 9, 2026, Apple unveiled the iPhone Duo, its first foldable phone, starting at $1,999 with US availability from October 23. The hinge and the titanium body are the headline. The more interesting story for anyone tracking AI is what is bolted underneath: a new A20 Pro chip whose Neural Engine now runs the phone's photography, its assistant and a chunk of its interface locally, while Apple's biggest AI moves this year have happened somewhere else entirely, on servers it does not fully build itself.

$1,999iPhone Duo starting price256GB, US
7.6inInner display50% larger than iPhone 18 Pro Max
5.4inOuter display90% of iPhone 18 Pro screen area
Dual 16-coreNeural Engine cores2x A19 Pro compute power
31h / 44hVideo playbackInner vs. outer display
Oct 23US availabilityPreorders from Oct 16

What Apple announced

Opened flat, iPhone Duo carries a 7.6-inch inner display and is, Apple says, the thinnest iPhone it has ever shipped. Closed, the 5.4-inch outer display covers roughly 90 percent of the screen area on the iPhone 18 Pro, close enough that Apple is positioning it as a phone you do not have to unfold for most tasks. The body is grade 5 titanium with a mirror-polished finish, and pricing starts at $1,999 for 256GB, rising through 512GB, 1TB and 2TB.

Battery life is split by which display is doing the work: up to 31 hours of video playback on the inner screen, up to 44 hours on the outer one, and about 24 hours when both are used roughly equally. Apple credits a custom vapor chamber and a dual-battery architecture for holding that up while running a noticeably larger, brighter inner panel.

The chip doing the AI work

The A20 Pro is the same chip Apple put in the iPhone 18 Pro this cycle, and its Neural Engine is the part that matters for everything that follows. Apple describes it as a "dual 16-core" Neural Engine delivering twice the compute power of the A19 Pro's, alongside a 6-core CPU it says is up to 20 percent faster generation over generation. That doubling did not come out of nowhere. Apple's Neural Engine has been scaling for almost a decade, from 2 cores in the A11 Bionic in 2017 to 16 cores by the A14 Bionic in 2020, where it has sat, generation after generation, until this year.

A decade of Neural Engine cores

Apple-specified core counts at launch, A11 Bionic (2017) through A20 Pro (2026). Apple describes A20 Pro's Neural Engine as "dual 16-core," with 2x the compute power of A19 Pro. Core count alone does not capture throughput; Apple has also raised operations per second within a fixed core count across generations.

Every generation since A14 has shipped Apple's largest on-device AI accelerator yet, and the phone in your pocket has run one for almost a decade.

A bigger, faster Neural Engine is what lets Apple run real-time inference on camera and audio streams without waiting on a network. It also needs somewhere to dump the heat that sustained inference generates, which is the practical reason a "thinnest iPhone ever" ships with a vapor chamber and a split battery: running a model continuously, not just answering one prompt, is a thermal problem as much as a compute one.

What "on-device AI" means on this phone, concretely

The most visible new feature is Smart Take: an on-device model, running on the A20 Pro, watches the camera feed in real time and fires the shutter automatically once everyone in a group photo is actually looking at the lens and not mid-blink. It is paired with a 48-megapixel fusion main camera and a 48-megapixel ultrawide with 2x optical zoom, and Apple is also shipping audio tuning that adapts to how the phone is physically positioned, plus an animated outer-display mode aimed at the person being photographed.

John Ternus, who became Apple CEO on September 1, 2026, photographed at an Apple event in March 2026
John Ternus, Apple's CEO since September 1, 2026, photographed at Apple's 50th-anniversary event at the Grand Central Terminal store in New York in March 2026, before the iPhone Duo announcement. Photo by Tessa Bury via Wikimedia Commons, CC BY 4.0.

Every one of those claims is Apple's own, made on stage and in its newsroom post, and none of it has been independently benchmarked yet. What is notable is the shape of the claim: it is a narrow, well-defined task (detect open eyes and a forward gaze, then time a shutter press) running continuously and locally, which is exactly the category of work a Neural Engine is built for.

Why the whole phone does not just run on-device

Apple's own research describes its on-device language model, the latest generation called AFM 3 Core, as roughly 3 billion parameters, small enough to fit and run inside a phone's power and memory budget. Its cloud-side sibling, AFM 3 Core Advanced, is larger and sparse, activating only 1 to 4 billion of roughly 20 billion parameters per request. Both are tiny next to the models the rest of the AI industry now calls frontier-scale, and that gap is not an accident. It is the ceiling that on-device inference runs into.

Two places to run inference

Running a model on the phone's own Neural Engine and running it on a data center server trade off along the same handful of axes. Apple Intelligence uses both, routing by task.

On-device (Neural Engine)
Cloud (data center)
Latency
On-device:No network round trip; response starts as soon as the chip finishes.
Cloud:Bounded by network conditions plus queueing at the provider.
Privacy
On-device:Personal data (messages, photos, on-screen context) never leaves the phone.
Cloud:Data leaves the device; Apple's answer is encrypted, non-retained processing inside Private Cloud Compute.
Cost
On-device:No per-query server bill; the cost is paid once, in the silicon.
Cloud:Every query consumes data center compute and power, billed per request or subsidized by subscription.
Model-size ceiling
On-device:Bounded by on-package memory and power draw; realistically billions, not trillions, of parameters.
Cloud:No practical ceiling; frontier models run to hundreds of billions or trillions of parameters.
Availability
On-device:Works with no signal, on a plane or underground.
Cloud:Requires a working network connection to the provider.

A 16-core Neural Engine is fast enough for narrow, well-scoped tasks. It is not a substitute for a data center when a query needs frontier-scale reasoning.

Latency, privacy and availability favor keeping inference on the phone. Model size and raw capability favor sending the hard problems to a data center. Apple's answer has been to do both, and to try to keep the privacy properties of on-device processing even when the work leaves the device, by routing cloud requests through what it calls Private Cloud Compute, encrypted, non-retained processing on servers Apple controls rather than a general-purpose cloud.

“

A 3 billion parameter model on the phone and a licensed 1.2 trillion parameter model in Apple's own data centers are not two versions of the same idea. They are the two ends of the tradeoff Apple is trying to paper over with one brand name.

On the size gap inside Apple Intelligence

The hybrid answer, and why it needed a new CEO's first launch

Apple logoApple
Google logoGoogle
Anthropic logoAnthropic
OpenAI logoOpenAI
Apple's AI stack now spans its own on-device models, a licensed Google model for the hard cases, and an optional ChatGPT hand-off inside Siri.

Apple has never shipped its own frontier-scale model, and 2026 is the year that stopped being a secret plan and became a public architecture. According to reporting from Bloomberg in late 2025, confirmed by Apple's own January 2026 announcement, Apple is paying Google roughly $1 billion a year to license a custom Gemini model with about 1.2 trillion parameters, eight times the size of the 150-billion-parameter model Apple's cloud features previously used, to power the more demanding parts of the new Siri. Apple reportedly also evaluated Anthropic, whose bid came in around $1.5 billion a year. The custom Gemini model runs inside Apple's Private Cloud Compute rather than on Google's own infrastructure, and Apple has said user data is not shared back to Google. Separately, Siri has offered an optional hand-off to OpenAI's ChatGPT for questions Apple's own model declines to answer, an older, narrower integration than the new Gemini-powered core.

That architecture shipped its first marquee hardware on September 9 under John Ternus, who became Apple's CEO on September 1, 2026, after 15 years under Tim Cook, who moved to executive chairman. Ternus spent his Apple career in hardware engineering, and the iPhone Duo, a hinge, a chip and a battery problem before it is an AI problem, is a natural first product for him to stand behind. The AI strategy underneath it, small local models plus a licensed outside frontier model, was largely set before he took the job.

Siri AI's rollout is staggered, and not everywhere at once

The redesigned assistant, Siri AI, launches in English with iOS 27 on September 14, with French, Japanese, Korean, Portuguese and Spanish following in October. Apple says it can now draw on personal context from messages, email and photos, read what is on screen, and take actions across apps rather than only answering questions. In the European Union, Apple is reportedly withholding the assistant from iPhone, iPad and Apple Watch at launch while allowing it on Mac and Vision Pro, a split that points at unresolved regulatory questions around exactly the personal-context features that make the new Siri more capable than the old one.

What this means

Strip away the hinge, and the iPhone Duo is Apple's clearest public statement yet of how it thinks AI should be split between a phone and a data center: small, fast, private models for continuous, narrow tasks on the device itself, and a licensed frontier model, run on Apple's own servers, for the questions that need real scale. None of the individual pieces is unique to Apple. What is distinctive is doing it without a frontier model of its own, buying that piece from a company, Google, that is also one of the more direct competitors to the assistant it is now powering.

That same lesson, that no single model is the right one for every task, is the reason it is worth being able to reach more than one of them from the same place. Metir gives you Gemini, GPT, Claude and Grok models side by side in one workspace, so switching between a fast local-feeling answer and a heavier reasoning pass does not mean switching apps.

Sources:

  • Apple unveils iPhone Duo - Apple Newsroom
  • Apple announces its foldable iPhone Duo - MacRumors
  • Apple unveils its first foldable, the iPhone Duo - TechCrunch
  • Siri AI launches in English this month, will add five languages in October - 9to5Mac
  • Tim Cook to become Apple Executive Chairman, John Ternus to become Apple CEO - Apple Newsroom
  • John Ternus succeeds Tim Cook as Apple CEO after 15 years - Al Jazeera
  • Apple Plans to Use 1.2 Trillion Parameter Google Gemini Model to Power New Siri - Bloomberg
  • Apple picks Google's Gemini to run AI-powered Siri coming this year - CNBC
  • Apple reveals new AI architecture built around Google Gemini models - MacRumors
  • Introducing the Third Generation of Apple's Foundation Models - Apple Machine Learning Research

Image credits

Header image: aerial view of Apple Park, Apple's Cupertino, California headquarters, June 2024, by Nils Huenerfuerst via Wikimedia Commons, CC0. In-body photo of John Ternus, Apple CEO since September 1, 2026, photographed at Apple's 50th-anniversary event in March 2026, by Tessa Bury via Wikimedia Commons, licensed under CC BY 4.0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational