It stops when you close the lid
The assistant you pay for lives in a tab. Shut the laptop and the research job dies, the inbox goes untriaged, the scheduled thing never runs. You are paying for a capability that is only awake when you are.
In stock · 2–5 working days
Always on, in your room. Bring the Claude, GPT or Gemini key you already pay for — no subscription to us, and not a cent on top of your tokens.
Compatibility
No logo wall here — a row of enterprise badges nobody can check is the fastest way to lose a reader who tries. What is true instead: the account you already pay for, the weights you already run and the tools already on your machine all point at it on the first evening.
Why a box
The models are already good. The keys are already in your password manager. What is missing is a machine that stays awake and does something with them.
The assistant you pay for lives in a tab. Shut the laptop and the research job dies, the inbox goes untriaged, the scheduled thing never runs. You are paying for a capability that is only awake when you are.
A Claude plan. A GPT plan. Then a third invoice from a service that resells you the same tokens at a markup, in its own credits, on its own servers — for the privilege of running them while you sleep.
Contracts, patient notes, unreleased code. So you rewrite the real question into a vague one, get a vague answer back, and quietly do that half by hand instead.
Bring your own key
Paste in the provider you are already billed by. The box talks to Anthropic or OpenAI exactly the way your laptop does — we run no proxy, take no cut and never see a token.
Anthropic
Claude
Your API key or Claude plan. Billed by Anthropic, not us.
OpenAI
GPT
Drop in the key you already use for everything else.
Gemini
Long-context work, on the account you already have.
OpenRouter
300+ models
One key, the whole catalogue, switchable per task.
Ollama
Local
Runs on the box itself. No key, no bill, nothing leaves.
When nothing can leave
There is a model living on the drive, running on the box’s own silicon. No key, no invoice, no packet off your network — so the contract, the patient note and the unreleased repo can finally be the actual question instead of a vaguer one you rewrote to feel safe.
You choose, per task
A job arrives
Stays here
The model on the box
No key, no bill, nothing off the network.
Goes out
The key you already hold
Billed by Anthropic or OpenAI, exactly as now.
Already on the drive
What it can hold
14 tok/s
A 8B-class model at 4-bit, fully resident in 8 GB of shared memory. It is not a frontier model and we will not pretend otherwise — it is the one that answers when the answer is not allowed to leave the room.
How you reach it
hermes.local
An OpenAI-compatible endpoint on your own network. Point your editor, your shortcuts or a six-line script at it and the only thing that changes is the base URL.
Not only chat
Speech and search
Whisper transcribes and Nomic embeds, both on the box. The recording of the call and the folder you searched never become anyone else's problem.

While you sleep
Because it isn’t your laptop. Long jobs finish overnight, the inbox is triaged before you open it, and it messages you when something needs you — not the other way round.
The 90-Second Setup
No Docker compose file, no CUDA versions, no quantisation rabbit hole. The hard part is already done and soldered down.
Power and Ethernet. It boots, finds itself on your network, and answers to hermes.local. No account to create, no setup wizard, no waiting list.
Anthropic, OpenAI, Google, OpenRouter — whichever you're already paying for. The box uses your key and your plan; your provider bills you exactly what they billed you yesterday. We never see a token.
Or the web UI, or the OpenAI-compatible endpoint your existing tools already speak. It's awake at 3am with your laptop shut, because it isn't your laptop.
What people point it at
None of this is a demo. It is the middle of the week, handled by something that does not need you awake for it.
triage, drafts, follow-ups that actually go out
Browser
fills forms, pulls data, clicks through the boring bit
Research
long jobs that run while the laptop is shut
Reminders
it messages you, not the other way round
Dev work
reviews, refactors, scripted chores on a schedule
Not this
training models, video work, or anything that needs more than 8 GB at once
In a real room

An owner’s desk. One cable, about 20 W, and a machine that is still working after you have walked away from it.
Agent Skills
Open, published skills — reading spreadsheets, filling PDFs, driving a browser, writing decks. One line each, and it goes and gets them itself.
$ hermes skills add pdf
installed · ready in this conversation

Built, not rendered
Specifications
It draws about twenty watts and sits there doing your work while the laptop is shut.
NVIDIA Jetson Orin Nano Super
LPDDR5, shared with the GPU
NVMe SSD
15–20 W typical — a fifth of a lightbulb
4-bit quantised, fully resident
7B at Q4_K_M, single stream
plus llama.cpp and an OpenAI-style API
single managed fan, not silent
The trade
Compared against a hosted agent service, not against the model providers themselves — you keep those. HermesBox is where they run, not a replacement for them.
The HermesBox Promise
Buying hardware from a company you had not heard of this morning is a leap. These are the terms that make it a small one.
Plug it in, connect your keys, push it hard for a month. If it isn't what you wanted, it goes back and we refund it. One email, no form, no restocking fee.
Your key, your provider, your invoice. We take no cut, add no markup and run no proxy — the box talks to Anthropic or OpenAI exactly the way your laptop already does.
The hardware is a one-time purchase and it works fully without paying us anything again. There's no tier that unlocks the thing you just bought.
There is no analytics to switch off, because there is none to begin with. The box has no reason to report on you and it doesn't.
Questions
No, and it doesn't have to be. The box runs 7–8B models locally, which handle the everyday middle of the work — summarising, extraction, classification, drafting, first-pass code. For anything harder you point it at Claude, GPT or Gemini on your own key and get exactly the frontier model you're already paying for. The local side isn't the ceiling, it's the option for documents you'd never upload.
No, and you can't. You connect your own Anthropic, OpenAI, Google or OpenRouter key and the box calls those providers directly. They bill you at their rates, on the account you already have. We take no cut, add no markup and run no proxy in between — the €549 is the entire commercial relationship.
That's the intended way to run it. Paste in the keys you already have and the box uses your existing plans — nothing is duplicated and nothing new starts billing. If you'd rather not use a cloud provider at all, the local models on the box work with no key and no bill.
Because the dev kit isn't the product. That figure is a US list price before VAT, shipping and import — landed in Europe it is already meaningfully more — and what arrives is a bare board with an empty M.2 slot. Getting from there to this adds the 512 GB NVMe drive, the case, the active cooling, the assembly and the burn-in test, and the work of having the agent, the runtimes and the local weights already talking to each other on first boot. The price then includes EU VAT, delivery, thirty days to send it back for a full refund, and the 2-year legal guarantee of conformity. There is a margin on top of that — we build these, support them and answer the email — and a breakdown engineered to land on exactly €549 would be an insult to anyone able to do the sum in the first place. If it still doesn't look worth it, the thirty days are there so you can find out with our money rather than our argument.
Don't, yet. Trust the terms instead: use it for thirty days and send it back if it isn't right. And note what you are not being asked to trust us with — your API keys stay on hardware in your room, your tokens are billed by your provider, and there's no subscription that can be raised on you later.
No. It arrives with Ollama and llama.cpp installed, weights already on the drive, and a web UI at hermes.local. If you do want to go deeper — pulling different models, adjusting context, exposing it over Tailscale — that's all there, just not in your way.
Yes. It exposes an OpenAI-compatible endpoint on your local network, so anything that lets you set a base URL works: Cursor, Continue, Zed, Raycast, Obsidian plugins, n8n, LangChain, LlamaIndex, or a five-line Python script. Day to day, most people just message it on Telegram.
On the box, on your network. They are not synced to us, not held in a hosted dashboard, and not readable by anyone who hasn't got the machine. That's the practical difference between this and an agent service you sign into: revoking our access means unplugging a cable.
They cost nothing — they aren't ours to sell. A skill is a small folder of Markdown that teaches an agent one job properly: filling a PDF, cleaning a spreadsheet, reviewing a diff, writing to a brand guide. The box installs any of the 652 in the open Agent Skills ecosystem in a single line, and you can read one before installing it and edit it after. Your own folder of instructions works exactly the same way. Skills run against whichever provider you pointed the box at, so the one that reads a contract can stay on the local model while the rest go to Claude or GPT on your key.
Yes — anything Ollama or llama.cpp will run that fits in 8 GB, which in practice means GGUF up to about 9B at four-bit. The models it ships with are a starting point, not a walled garden. Pull your own with one command.
No. It has a single managed fan and it is quiet, but a page that told you a computer with active cooling makes no noise would be lying to you about the first thing you'd notice.
It ships from Bulgaria to the EU, UK, Switzerland and Norway. Shipping and EU VAT are included in the price shown — the number at checkout is the number you pay.
Send it back within thirty days of it arriving and we refund it in full. One email — no form, no restocking fee, no explanation required.
Still unsure? Ask us directly — a person answers.
The offer

Assembled, loaded and left running before it is boxed. Bring the key you already pay for and it works on the first boot.
EU VAT and delivery included — the number at checkout is the number you pay. Ships from Bulgaria to the EU, UK, Switzerland and Norway. 2-year legal guarantee of conformity.
In the box
Not ready yet
Get the setup guide — which providers plug in where, what an 8 GB local model actually handles, and a monthly note on what people are automating with theirs.
One email a month at most. No tracking pixels, and unsubscribe from any of them.