Fine-tuning Kev 9B past Kev 27B and Jev
Kev 1.0 shipped yesterday with open weights and a fine-tune recipe. the recipe is the easy part. the hard part is a few thousand labelled requests in your app's own shape. this is how I got those and trained Kev 9B with dopp, start to finish, for about a dollar of GPU.
the eval
Kev's repo has a frozen eval called transfer-v4: 764 records from datasets Kev never trained on. one part of it is 80 short sentences from dair-ai/emotion, each labelled sadness, joy, love, anger, fear or surprise. Kev's own numbers on those 80: Kev 27B 58.75%, Jev 58.75%, Kev 9B 57.5%.
first I checked our hosting: I ran the whole transfer-v4 set through dopp to Kev 27B with Kev's own benchmark script (kev.benchmark --remote). it scored 0.8506, the same as Kev's published 0.851, and emotion came out at 47 of 80, the same 58.75%. so the model behind dopp is the one Kev measured.
you need: a free dopp.sh account, or the repo on your own machine with a Modal account for the GPU part.
1. a route that answers with Kev 27B
Routes → New route, name it, and pick Kev 27B (open) as what answers. you get a key. your app sends the same request it sends Jev today (state plus typed questions) to https://dopp.sh/v1/systemone with that key, and Kev 27B answers. every request is kept, with who answered and how long it took.
the first request after a quiet spell wakes the GPU, which takes about a minute. after that it answered in about 0.6 to 0.8 s end to end for me.
2. labelled rows
a fine-tune needs the right answer for each request. in your own app that's your requests plus a label: you fix answers on the Requests page, or you pick who labels them (a person, an LLM, a dataset).
for this test I wanted rows like the eval's, so I used Add requests → Find datasets, typed dair-ai/emotion, picked the train split and labels from column label. the dataset's own labels become each row's label, and rows that bring a label don't get sent to any model. 2,500 rows came in, and dopp set 1 in 10 aside to check each version. the 80 eval rows come from the dataset's test split, so none of them are in training.
3. train Kev 9B
Models → Train a model → base Kev 9B, 2 passes, labels from each request. it trained on 2,531 rows on an H100 at about 0.11 s per row per pass, so roughly 9 minutes of training plus loading. dopp's own price for the run was $2.42; the GPU underneath cost about half that.
the new version, emotions v1, agreed with the labels on 84.9% of the 279 rows set aside.
4. switch the route and measure
on the route, Setup → Start from → Your model only → Save. the app keeps sending the same requests; now emotions v1 answers them. I sent the 80 eval rows through the route like app traffic:
| on Kev's 80 emotion rows | correct |
|---|---|
| Jev | 47 (58.75%) |
| Kev 27B | 47 (58.75%) |
| Kev 9B, released | 46 (57.5%) |
| Kev 9B, fine-tuned through dopp | 60 (75%) |
80 rows is small and the emotion labels are noisy (they came from hashtags), so take the exact number loosely. the direction is clear though, and it's the gap Kev's README talks about: a general model on a task that isn't quite its own.
what it costs
on dopp.sh: the training run above was $2.42. hosting a model is $29 a month, or it answers from your credit by the hour while it's awake. self-hosted, it's the repo plus your own Modal account (Kev 9B trains on one H100).
run it yourself
everything above is in the repo: the proxy, the dashboard, Find datasets, the trainers. Kev 1.0 runs through the same engine as its own kev-finetune skill: a delta from the released checkpoint with its own recipe.