Can you run Jev offline?
Not Jev itself: TypeSafe doesn't release its weights. What you can do is train a small model on the requests your app already sends to Jev, and run that model on your own machine with the network off. Below is the whole thing, start to finish, with a Home Assistant voice assistant that ends up switching lights with the router unplugged.
The setup
HA-Jev is a Home Assistant integration that turns every voice command into one Jev request. The request carries the command, a snapshot of your devices, and about a dozen typed questions: what should happen (turn_on, turn_off, set_brightness…), which device, is there a delay in it, does it mention brightness, and so on. Jev answers, Home Assistant acts.
It works well. It also sends every "turn off the kitchen lights" over the internet, and its own README measures 250 to 580 ms warm and 700 to 900 ms after idling. If the internet is down, so is your voice control. That's the thing people keep asking about on the Home Assistant forum and on Hacker News: can this run locally?
1. Change one URL
HA-Jev has an address field under Advanced. Put https://dopp.sh there and a Dopp route key in the key field. Nothing else changes: Jev still answers every command, and Dopp keeps each request.

Say a few commands and they show up in the request log, with the command up front and the house snapshot folded to "entities: 33 items".

2. Is it enough yet?
The route page counts labelled examples for every answer of every question and colours them: green at 20 or more, amber from 5, red under 5. Answers your requests never produce are folded away instead of shown as failures.
After a handful of real commands, almost everything is red. That's expected: nobody says "turn off the living room fan" twenty times on day one.

3. Fill the gaps
Two ways, both on the same page.
A public dataset. Find datasets takes a name. Amazon's MASSIVE (SetFit/amazon_massive_intent_en-US) is 11,514 English voice-assistant commands written by people. I added {{MASSIVE_ROWS}} of them. Most are about alarms, music and weather, not devices, so they mostly teach the model what "not about the house" looks like. Useful, but the device actions stay thin.
Generate. Each red answer has a "Write 20" button. It writes new commands aimed at that answer, in your house's terms ("shut the kitchen window", "switch off the ceiling lights", "close the pergola roof"), and drops each one into a real request so the device list and questions stay exactly as your app sends them. {{GENERATED}} commands later the actions were green.

4. Train the small one
"Which base to train" has a tab for where the model will run. On "Downloaded, offline" the pick is Tiny: a 34 MB model (bge-small, int8) with one output per question. It reads the request's text field, here the command, and knows the questions and options it was trained on.
Training took {{TRAIN_SECONDS}} seconds and cost {{TRAIN_COST}}.
The Models page shows held-out agreement: how often the model gives the label on requests it never trained on. {{HOLDOUT_OVERALL}} overall, but the overall number is carried by the easy yes/no questions, so it also shows each question on its own: {{HOLDOUT_BY_QUESTION}}.
5. Download it and unplug
"Download to run offline" gives you a folder: the model, its tokenizer, a README that says what it knows and how well, and serve.py, a small server that answers the same POST /v1/systemone requests Jev does.
pip install onnxruntime tokenizers numpy
python serve.py --host 0.0.0.0
Then the second URL change: HA-Jev → Reconfigure → address http://<that machine>:8787.

"turn off the kitchen lights": {{OFFLINE_RESULT}}. Answers take about {{OFFLINE_MS}} ms, and nothing leaves the network.
It isn't perfect. "switch on the ceiling lights" came back unsure, and HA-Jev handed it to Home Assistant's own agent instead of guessing. That's the right behaviour, and the fix is the same as before: a few more examples of that phrasing, one more training.
Will it run on a Raspberry Pi?
It should. The folder is Python plus onnxruntime, both of which install on 64-bit Raspberry Pi OS, and I ran it in an ARM64 Linux container. I haven't timed it on a Pi yet. At 34 MB and about 5 ms per command on a laptop CPU, there's a lot of room.
When not to do this
If what you're doing is search, like Linear's emoji picker picking from thousands of emoji, a classifier is the wrong tool: there's no fixed set of answers to learn. Max Leiter's embedding approach is the better fit there. This works when your app asks the same questions with the same options over and over.
Questions
FAQ
Can I fine-tune Jev itself?
No. TypeSafe doesn't release Jev's weights. You train an open model on the same requests.
Do I have to stop using Jev?
No. Keep Jev answering while you collect, and switch when your own model is good enough, per route.
How much did this cost?
{{TOTAL_COST}} in all, on the free credit every account starts with.
What about Laya or bigger models?
They're on the same "Which base to train" list. Laya reads the questions at run time and handles new options without retraining, but it's 842 MB. For a device in a cupboard, Tiny is the one.
What Dopp is. Dopp is a drop-in proxy for any Jev-compatible decision API: TypeSafe's Jev, Kev, GLiNER, simple-jev, or your own endpoint that answers typed questions about a state. You change one URL; every request still gets your current API's answer and is recorded as training data. When you have enough, one button trains a small open model on your own requests, measures it against the API it learned from on requests it never saw, and lets you choose who answers: the API, your model with the API as backup, or your model alone, on our servers, or offline, even in your users' browsers.