What I learned trying to use Jevto judge AI-generated photos
Jev can't see images, so its photo scores are only as good as the model describing each photo. What six describers taught me.
Jev, TypeSafe AI's new decision model, can't see images. For a photo-judging job, that means its score is only as good as the model that describes each photo to it, and the big model stays in the loop.
The hype, and my own fun with it
The model is Jev, from TypeSafe AI, and it's been all over my feed since its September 15 launch. You give it text and a list of questions, and it hands back probabilities in under a second, for a sliver of what a big model costs. On my Mac, calls came back in 505 to 776 milliseconds.
One of TypeSafe's launch demos has it playing Doom. Stefan Mai wrote that his team moved three Claude Haiku tasks to it and matched Haiku's accuracy at about a tenth of the cost.
I'd had my fun with it too. I heard about it on a podcast on the 18th and had it running before midnight. My second test was a made-up customer email: is this about shipping (1.0), does it need a reply (0.84), in 594 milliseconds.
A few days later it scored 23 job applications in one batch. All 23 landed under a cutoff borrowed from my interview scorer, a 3.0 threshold built for full interview transcripts. A few lines of text can't clear a bar set for a whole interview. That was a useful early lesson on its own: a cutoff that works on one kind of input does not carry over to a thinner one.
The job I cared about
On Sunday I gave it a job I cared about. I run pFrame, which makes product photos with AI, and I wanted a fast way to sort the good images from the ones with AI tells, like melted lettering on a label.
Just before midnight the plan came back with about 75 minutes of Claude vision work for 518 photos. I typed: "why do we need claude vision work on the 518 photos? i thought we're using jev on them?"
Jev can't see. Claude looks at each photo and writes notes, and Jev scores the notes. TypeSafe's docs say it plainly: pre-process images, audio and video into text or structured fields before you send them.
Then how does it play Doom?
The demo runs on game state written out as text. The launch post says so directly: the Doom demo is "not on images (yet...)". The same post is honest that a non-AI Doom bot could play better.
TypeSafe does have one image-capable product. Jev Router, listed on OpenRouter on September 25, accepts images. Its listing describes its job as picking the best model for each request.
The describer decides the answer
So Jev's score on a photo can only be as good as the notes it's handed. We tested six models as the note-writer, on 18 real photos and six AI images with their file info stripped, so the only way to catch them was to look.
Every run used the same Jev model and the same rubric. Same Jev, same cutoff, and the best caught six times as many AI images as the worst:
- Claude Opus: 6 of 6.
- DeepSeek V4.1 Flash: 6 of 6, but it also marked down three real photos.
- GLM-5.3-Flash: 5 of 6.
- Kimi K3: 4 of 6, and it marked down one real photo.
- Qwen3.8 27B, on my own Mac: 3 of 6.
- Gemini 3.8 Flash: 1 of 6. It was also the fastest, at about 11 seconds a photo.
The cutoff was 0.8. When I moved the cutoff afterwards, the free local Qwen sorted all 24 correctly: its lowest real photo scored 0.878 and its highest AI image 0.84, so a cutoff of 0.85 splits them cleanly. Six AI images is a tiny test, and I picked that cutoff after seeing the data, so I don't trust that yet. It is still the most interesting result in the set, because a free model on my own machine would change the cost picture completely if it held up on a bigger test.
Jev was mostly re-grading Claude's call
Claude's notes include its own read on whether a photo is real, so on the question I cared about most, Jev was mostly re-grading Claude's call. I haven't measured whether it beats Claude's own verdict.
I kept Claude as the eyes. The first real batch, 205 generated product scenes, took 28 minutes and about three cents of Jev, and most of those minutes were Claude looking. By my count, Jev's calls added up to roughly 12 percent of the worker time. Jev costs about a hundredth of a cent a photo.
It works. It just isn't what I thought I was getting: a way to take the big model out of the loop.
What I'd tell a business trying a new AI tool
Try new tools on your own work, since that's the benchmark that counts. A launch demo shows what a tool can do on the input its makers chose. Your job decides what the input actually is.
Three questions I now ask first:
- What does it take as input? If your data is images, audio or video and the tool reads text, something else has to do the looking, and you are paying for that too.
- What does the step before it cost? Jev's own cost was a rounding error. The describer was the time and the money.
- Where does your cutoff come from? My job-application cutoff came from interviews, and my photo cutoff treated six describers the same when at least one needed its own.
For photos, Jev made the scoring cheap. The looking still makes the call.
What have you tried recently that didn't live up to the hype?
What I changed in my AI workflows after trying Claude Opus 5.5
What changed when I audited my coding-agent skills with Opus 5.5: fewer procedural rules, clearer ownership, and stronger verification.
ReadHow I Automated Myself Out of a Job (In a Good Way)
I built an AI webmaster so my client could stop waiting on me. Here's how it works, what broke, and what I'd do again.
ReadWhy So Much AI Writing Sounds the Same, and What I Am Doing About Mine
A model can copy how I sound. It cannot decide what I think, which is why I'm spending more time writing, not less.
ReadWorking on something like this?
Bring the app or the process to a free 15-minute call. I will tell you what I would look at first, and whether I am the right person for it.