metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
On-Device AI
AI Agents
Funding
Apple Silicon
Personal AI

Underdog On-Device AI Assistant: Conway Backed by a16z

Conway Research launched Underdog, a free on-device AI assistant backed by a16z and Khosla. What the 27B model, Husky engine and benchmarks do and do not show.

Metir AI TeamOctober 7, 20268 min read
Underdog On-Device AI Assistant: Conway Backed by a16z

Conway Research, founded by 2025 Thiel Fellow Sigil Wen, has launched Underdog, an on-device AI assistant that the company says is free, ad-free and runs entirely on the user's own computer. Backers announced the funding on October 2, 2026, and TechCrunch covered the product launch on October 6. The Underdog on-device AI assistant arrives in a week when cloud-hosted personal agents from Meta and Instinct are drawing most of the attention and capital, which makes its opposite design choice the interesting part of the story.

Apple logoApple
Qwen logoQwen
Anthropic logoAnthropic
Meta logoMeta
Companies and model families referenced in this story: Apple silicon hardware, Alibaba's Qwen base model, Anthropic's Claude Opus 4.6 as a benchmark reference, and Meta's Muse as a rival.
27BParameters in the current modelfine-tuned from Qwen3.8-27B
4BParameters in Woofthe small model, under 2.5 GB
730Peak tokens per secondWoof on an M5 Max, company test
4.5xClaimed top speedup over MLXbest case, not typical
$0Price to the userfee on Stripe payments instead

What Conway announced

According to TechCrunch, Underdog launched as an invite-only beta. It currently uses a 27-billion-parameter reasoning model fine-tuned from Qwen3.8-27B and is powered by Husky, an inference engine Wen built to run models quickly on a user's own hardware. The assistant handles everyday tasks such as shopping research and homework help. Users are not charged and are not shown ads. Instead, TechCrunch reports, Underdog will take a small percentage of payment transactions the assistant makes through Stripe's payment rails, something like an interchange fee.

The funding was announced on October 2 in an a16z post. Andreessen Horowitz and Khosla Ventures are described as leads in TechCrunch's coverage, with Hummingbird, SV Angel and the Anthology Fund also named, along with angels including Stripe co-founder Patrick Collison. The round size and valuation were not disclosed in the coverage we reviewed, so this post does not estimate them.

One detail differs across reports. TechCrunch and Digital Today say the beta currently runs on Macs and Windows PCs, with Linux, iPhone and Android versions coming soon. Other outlets describe it as targeting Macs and iPhones. We follow TechCrunch's wording, which is the most specific: iPhone is planned, not shipping.

“

You don't need to sacrifice your privacy for the capability because they're just as capable.

Sigil Wen, Conway Research founder, via TechCrunch

The benchmark claim, read carefully

The headline claim is that Underdog's model compares favorably with Anthropic's Claude Opus 4.6 on some benchmarks. TechCrunch frames this as roughly the top performance of about six months earlier. Three qualifiers matter here:

  • It is the company's claim. We did not find an independent evaluation of the 27B model in the coverage reviewed.
  • It is "some benchmarks," not all. The reporting says the comparison holds on certain tests, not across the board, and does not name them in the sources we could read.
  • The reference point is not the current frontier. Matching a model from earlier in the year on selected tests is a meaningful result for a model small enough to run locally, but it is a different statement from matching today's best hosted systems.

This is consistent with how the small-model gap has behaved for two years. Open and fine-tuned models in the tens of billions of parameters have repeatedly reached the level that frontier models held several months before, helped by better training data, distillation and reasoning-style post-training. Underdog builds on that pattern by starting from an existing open base, Alibaba's Qwen line, and tuning it, rather than training a model from scratch.

Husky, MLX and what the speed numbers mean

Conway publishes a benchmark page for Husky that compares it with Apple's MLX framework. The test used an Apple M5 Max with a 40-core GPU and 128 GB of memory, and the model was Woof, Conway's 4-bit, 4-billion-parameter model. Conway describes sixteen prompt types, three runs per engine, and alternating run order.

Decode speed on a 4B model: Husky vs Apple MLX

Tokens per second, Woof (4-bit, 4B parameters), Apple M5 Max. Company-run benchmark, four of sixteen prompt types.

Apple MLXHuskyHusky + Flash
Function edit
163
614
730
Write function
155
170
545
Fix typos
159
463
535
Call summary
162
166
211

Source: Conway Research, Husky benchmark page (husky.underdog.ai). Not independently verified.

The "up to 4.5x" figure comes from one task, a function edit, with a speculative-decoding style option Conway calls Flash enabled: 730 tokens per second against 163 for MLX. On other tasks the gains are smaller. Writing a function reached 545 against 155 with Flash (3.5x), while a call-summary task moved from 162 to 211 (1.3x), and without Flash the summary task was essentially unchanged at 166. Conway also reports first-token latency of 29 to 39 ms for Husky against 137 to 151 ms for MLX on cached continuations.

Two cautions follow. First, these are results for the 4B model on a top-end laptop chip, not for the 27B model that the Opus comparison refers to. Second, gains were largest on tasks where output overlaps heavily with the input, such as editing and typo fixing, which is where speculative techniques tend to shine. Conway's own figures show the spread, which is more informative than the peak.

Close-up of the Apple M1 chip package with its two unified memory dies on a green circuit board
Apple's first M1 system-in-package, with the processor at left and two memory dies at right. It illustrates Apple silicon's unified-memory design and is not Underdog's test hardware, which was an M5 Max. Photo by Henriok, CC0.

Why Apple silicon memory sets the ceiling

On a Mac or iPhone, the CPU, GPU and Neural Engine share one pool of unified memory, which is why Apple hardware has become a popular place to run local models. It also defines the limit: the model has to fit in that pool alongside the operating system, the user's apps and the model's working memory (the KV cache that grows with conversation length).

Simple arithmetic shows the constraint. Storing 27 billion parameters at 4 bits each takes about 13.5 GB for the weights alone (our calculation, before cache and overhead). That is comfortable on a high-memory Mac and tight on a base laptop. A 4B model at 4 bits fits in under 2.5 GB, which is how Conway describes Woof, and that size is plausible on a phone. This is the practical logic of shipping two sizes: a small model for fast, frequent work and a larger one where the hardware allows. The sources we could read did not state how much memory the 27B model needs in practice, and we could not open the model's repository page, so we leave that figure out.

An iPhone 16 Pro Max in Desert Titanium resting on a MacBook
An iPhone and a MacBook, the two Apple device types a local assistant like Underdog is aimed at. The photo is generic and does not show Underdog running. Photo by Padgriffin, CC BY 4.0.

On-device versus cloud: the tradeoffs

The design choice carries consequences on four axes.

  • Privacy. A model running wholly on-device means, as TechCrunch puts it, that user data stays on devices the user already owns. TechCrunch also notes Underdog encrypts the keys to email and other accounts users authorize. Actions that reach outside, such as payments or sending email, still touch third parties by definition. Local inference removes the model provider from the data path; it does not remove the services the agent acts on.
  • Latency. No network round trip helps first-token time, and Conway's Husky figures point the same way. Hosted systems can offset this with far larger compute and faster batch serving.
  • Capability. Hosted frontier models still have more parameters and tools to draw on. Wen's argument is that for everyday tasks the gap no longer matters. That is plausible for shopping research and homework help, and untested in the sources we reviewed for long, multi-step work.
  • Cost. A local model shifts inference cost from the vendor's servers to the user's hardware and battery. That helps explain how a free, ad-free product can exist at all, with revenue coming from payment fees instead.

Personal agents compared: where they run and what they cost

As reported at launch. Glimmer is a model, not an assistant app.

Underdog (Conway Research)Local
Runs: On your Mac or PC (iPhone planned); invite-only beta
Price: Free, no ads; small cut of Stripe payments it makes
Muse (Meta)Cloud
Runs: Dedicated cloud VM per user (iOS, Android, web)
Price: Free for common use; paid plans for higher usage
InstinctCloud
Runs: Own phone number and computer for the agent; early access
Price: Not published
Muse Glimmer (Meta, open weights)Local
Runs: Model you run yourself on a single consumer GPU
Price: Apache 2.0 license, no per-token fee

Sources: TechCrunch, Meta, Metir AI coverage of Instinct and Muse Glimmer.

The personal-agent race

Underdog lands in a crowded field. Instinct, which we covered in our look at its $1 billion Series C, runs an agent with its own phone number and computer. Meta's Muse is hosted in a dedicated cloud VM per user, and the company has been extending it, as in our report on Muse for Small Business. Meta has also released an open-weight local model, covered in our Muse Glimmer post, which shows that even a cloud-first company sees a role for local models.

The rivals compete on convenience and reach. Underdog competes on where the data lives and on price. The business model is the least settled part. Taking a cut of payments assumes users will let the assistant buy things and that volume is large enough to matter, and the shopping-agent space already has friction, as the Amazon and Meta agentic commerce standoff showed.

What to watch next

  • Independent evaluation. Third-party results on the 27B model, including which benchmarks are behind the Opus 4.6 comparison.
  • The iPhone version. A phone release would test how much of the 27B experience fits within mobile memory and battery limits.
  • Real-world reliability. Whether an on-device model handles multi-step tasks with accounts and payments as dependably as hosted agents.
  • Unit economics. Whether payment fees can fund continued model work with no subscription.

For users who want to compare local and hosted models rather than commit to one, model-agnostic tools such as Metir let people use several providers side by side, which is one way to judge the privacy and capability tradeoff on their own tasks.

Bottom line

Underdog is a credible test of an old question in a new setting: how much of a personal assistant can run on hardware people already own. The verified facts are a funded startup, an invite-only beta, a 27B model built on Qwen, and company-published speed numbers on a smaller model. The unverified parts are the benchmark claims and the economics. Treat the Opus 4.6 comparison as a company claim until independent testing arrives.

Sources:

  • Silicon Valley's AI wunderkind launches Underdog | TechCrunch
  • Investing in Conway, the creator of Underdog | a16z
  • Husky benchmark page | Conway Research
  • a16z leads Conway Research's first round | Runtime Wire
  • Conway Research lands a16z backing for Underdog | Crypto Briefing
  • Underdog launches, revenue to come from payment fees | Digital Today
  • Conway Research describes Underdog's on-device AI and Apple benchmarks | TokenPost

Image credits

  • Header and first in-body image: Apple M1 chip package, photographed by Henriok via Wikimedia Commons, dedicated to the public domain under CC0. Reviewed October 7, 2026. It does not depict Underdog or an M5 Max.
  • Second in-body image: iPhone 16 Pro Max on a MacBook, by Padgriffin via Wikimedia Commons, licensed CC BY 4.0. Reviewed October 7, 2026. It does not depict Underdog.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational