iPhone app · release candidate in progress

Find the best fit for your real task.

Test your prompts across iPhone, cloud, and your own AI. Compare quality, speed, cost, and data path. Keep the winner—and the evidence.

$1.99 planned launch priceOne-time purchase · optional provider charges stay separate
BUNDLED PRACTICE · NOT LIVE
Real job

Summarize private meeting notes

2 targets

“Return three decisions, each owner, and the next deadline. Do not add facts.”

Apple Intelligence baselineYOUR WINNER
3/4 checks1.8s total$0 per runDevice path
Sample cloud modelFASTEST
4/4 checks1.1s total$0.004 est.Internet path
Why it won

The notes stay on the device. Privacy matters more than the extra quality point for this job.

One fixed promptEvery candidate gets the same job.
Human-selected winnerFastest does not mean best overall.
Saved decision recordRerun when a model, route, or device changes.
THE LOOP

From “which model?” to a decision you can defend.

Benchlet is not another generic chat screen. It helps you qualify a model for one job, under the constraints that actually matter.

01

Name the job

Start with a prompt you already care about. Fix the task, required format, privacy needs, available routes, and budget.

02

Run a fair test

Send the same prompt to at least two candidates. See exact routes, output limits, cost status, and what leaves the phone before cloud execution.

03

Keep the winner

Review instruction-following, usefulness, format, speed, estimated cost, and data path. Save your winner and the reason.

THREE HONEST ROUTES

Where the work runs matters as much as the model.

Apple baseline

Uses Apple’s system language model on a compatible iPhone. Benchlet does not claim an arbitrary Hub repository runs on-device.

On device

Hugging Face Cloud

Uses your inference token and the exact provider shown before each run. Credits, quotas, pricing, and provider policies may apply.

Internet

Your server

Connect an OpenAI-compatible endpoint on a Mac, PC, NVIDIA system, home server, or workstation that you control.

Your endpoint
EVIDENCE BEFORE HYPE

Know what is measured, estimated, claimed, or still unknown.

E

Exact metadata from the repository or live provider catalog.

M

Benchlet measurement from your recorded run environment.

Benchlet estimate that needs a real test before you rely on it.

?

Unknown when the available evidence cannot support a claim.

“The fastest model can still lose because it broke the required format, sent sensitive text online, or cost more than the job justified.”

Benchlet decision principle
BRING YOUR OWN ROUTES

No mandatory Benchlet account.

Saved models, cached briefs, comparisons, and credentials are designed to stay on your iPhone. Cloud prompts go to the service you explicitly choose. Server prompts go to the endpoint you configure.

Read the working privacy policy
PRELAUNCH

Use one real prompt. Test two routes. Make one decision.

The first validation group is limited to 20 qualified testers. TestFlight access opens only after the exact release candidate passes physical-iPhone smoke testing.

TestFlight list opening soonBenchlet is not yet available on the App Store. No purchase or signup is being accepted on this preview.