Everything else on this site is measured. This page is just my own setup, running on a Mac mini M4 with 24 GB, with the reasons and the bills attached. I rewrite it when something changes, and the date above is the last time I checked it myself.
AgentsMac mini M4, 24 GB
Hermes AgentNous Research
The harness everything else runs inside
The harness decides more than the model does. What the agent can see, when it stops, whether it asks before doing something you cannot undo. Hermes handles that part, and I can swap the model underneath without rebuilding anything, so the rest of this list can change while my setup stays put.
Daily driver$185.91 last week, 6B tokens
DeepSeek V4 FlashDeepSeek
Almost everything, almost all day
The fastest thing I have and still the cheapest by a wide margin. The price rise in August hurt, and it is still a good deal: six billion tokens last week cost me $185.91. That is the number I would want to see on a page like this, so there it is. I am also looking around, because somebody is probably serving the same model cheaper than I am paying.
Worth knowing. One thing to be ready for: DeepSeek has no vision at all. In a DeepSeek only setup you paste a screenshot, the model answers confidently, and it never saw the image. No error, no warning. I run a small vision model locally on the Mac to cover it, which took an evening to set up. Finding out the hard way costs a lot more, so if you go all in on DeepSeek, sort that out first.
What I switch to when a task is genuinely difficult
It holds up on the problems where the daily driver starts guessing. It also does not lecture me. No moral commentary on ordinary questions, which sounds like a small thing until a model has refused to do its job and you are arguing with it instead of working.
For working out what to build rather than building it, and for anything long enough that losing the thread halfway costs more than the model does. I tried its workflows recently and they are the best thing I have used in a while. On the subscription rather than the API, because paying per token for work this long is a number I would rather not watch.
The best memory I have used, and I have been on it a long time. It is open source, so it runs on my own machine rather than somebody else's.
Very fast, and it stores everything rather than a summary of everything.
Fits Hermes perfectly, which is why the two live on the same box.
One minute to install on a VPS if you would rather not run it locally.
Worth knowing. My honest take: since I started using it I have completely forgotten what context window compacting is. It stopped being a thing I think about. Whether the window is 250k or a million makes no difference to me any more, and that turned out to matter more than any benchmark score.
Web data
FirecrawlFirecrawl
How the agents read the web
Just good. It is fast, it gets me the page as usable text instead of a pile of markup, and I have never had to think about it twice. That is the whole review.
How it got here
Every change to this setup, newest first. A setup you can watch change is worth more than one that claims to be right.
Started keeping this page
Hermes on the Mac mini, DeepSeek V4 Flash for the day, Kimi K3 for the hard parts, Opus 5 for strategy. Everything after this is the record of what changed and why.
DeepSeek raised prices
Peak output on V4 Flash went from $0.28 to $1.32 per million tokens, half that off peak, announced three days earlier and live on 16 August. Still cheaper than everything else I could switch to, so it stays as the daily driver, but it is no longer free money and I started looking at other providers serving the same model.