The thinking, on your Mac.
Sting is our own engine for running AI on an ordinary 16 GB MacBook. For the two jobs Smart Jelly does most, finding what you owe and noticing when it’s done, it comes close to the cloud, in under a second.
Nearly as good as the cloud, on the jobs that matter.
Each job is asked exactly the way Smart Jelly asks it today. The score is how often the answer matches the cloud model the app uses. The black line is the cloud matched against itself: the most anyone can score.
Scored on test conversations no model learned from. Ranges and how we measured are in the research notes below.
Faster than asking the cloud.
No round trip to a server. Median times on a 16 GB MacBook with an M4 chip, against the app’s cloud call for the same job.
Most local AI runs any prompt. Sting runs one job.
It is written for one model, one kind of Mac and one task at a time. Knowing the job ahead of time is what makes a small model good enough.
Sting and Husky.
Husky, from Underdog, is the fastest engine for one model on Apple silicon that we know of. Husky is built for speed on any prompt. Sting is built for the right answer to one job, on the laptop people already own.
| Measure | Husky · M5 Max, 128 GB · their figures | Sting · M4, 16 GB · measured | What it means |
|---|---|---|---|
| Checking 8 drafted words, cost vs 1 | 1.5× | 1.03–1.1× | Sting checks a draft almost for free, so every accepted word is nearly pure gain. |
| A yes or no for a job | not offered | ~90 ms | Sting reads the answer in one pass; there is no text to write. |
| “Nothing to report” | writes the answer | one pass, then stops | Most chats hold no task; Sting stops as soon as it knows. |
| Scored on the app’s real jobs | not published | 81 / 92 vs cloud 93 / 99 | Every Sting number here is scored against the cloud model the app uses. |
| Words kept per read, copy-heavy job | 5.4 | 2.3–3.3 | Husky trains its drafter; Sting drafts from the conversation itself. |
| Plain writing speed vs MLX | up to 4.5× | ~1× (31 vs 33 words/s) | Both are limited by memory speed; Husky’s code gets closer to the limit. |
| Reading a new prompt | 5,150 tokens/s | 275–310 tokens/s | The M5 has matrix hardware the M4 lacks; Sting keeps job instructions in memory instead. |
Bold marks the side ahead on that row. Husky’s numbers are Underdog’s published figures on different hardware; Sting’s are measured on a 16 GB M4.
How we measured it.
What we’re working on.
On your Mac, in Smart Jelly.
Sting is coming to Smart Jelly as a third way to run it. For your to-dos, your conversations will be read on your Mac and nothing will be sent to a server. Everything else keeps working the way you choose today.
Download Smart Jelly today macOS 14+ · Apple silicon · On-device mode will need 16 GB of memory