I've found that using structured outputs solves this problem much better. Instead of letting a model generate only "A", "B" or "C" and looking at the probs, have it directly generate "Legitimate", "Spam" or "Phishing" or any other pre-defined option from a set of multi-token sequences. Behind the scenes it boils down to something quite similar, but you're not running into the risk that the model actually wanted to say "A phishing attempt seems likely, so answer (C) is correct.", which would lead "A" to have the highest probability in the first token. You can even use a reasoning budget this way either via inherent reasoning or a free-form part preceding the remaining output structure. You can also have it assign probabilities (either in words or numbers) using more complex output structures, but I would not rely on them much more than the token logprobs (they can still be quite good though).
I got this technique to work extremely reliably last year. However there were a bunch of caveats: 1) Firstly, you must institute a check that the multiple choice tokens dominate the output distribution. They should sum to 95% or more, ideally 99%, or the LLM is not following instructions properly. This is also the problem with constrained decoding - if the LLM really doesn't want to output a valid answer, the one you extract will not be high quality. 2) You need to ask it multiple times, permuting which option corresponds to which letter, and average the results. LLMs are surprisingly biased towards picking "A", especially if they're otherwise not sure. 3) For the same reason, performance improves if you frame the prompt as if it were the middle of a quiz. "Question 1" carries baggage that "Question 12" doesn't. 4) You must be exceedingly careful with tokenization.
But when all was said and done, I got a general purpose A/B classifier that gave high resolution quantitative output for the cost of a couple dozen tokens ingested and a couple inference passes.
I ran some tests using GPT-4 to do some basic classification a couple years ago. On ambiguous options which had to be escalated to a human, the LLM would regularly output something like a 99.8% probability, compared to 99.99% for a correct answer.
https://til.simonwillison.net/llms/llama-cpp-python-grammars
Prompt part: "What is better, toast or bread?"
Incomplete answer part: "The answer to this question is "
and then have the LLM finish the answer. I did this with subtitle translation using llama.cpp (with Python) and had great success. Just past 5 already translated subtitles as the incomplete answer, and the LLM infallibly just continues to translate. No markdown, and usually no talkback if the subtitles contain nasty subjects like bioweapons or nuclear stuff. It just works.
To completely squash the issue, a few cheap LoRa iterations will do the trick just fine.
I think we can all agree that Jev is not rocket science. It's a good idea executed well, with marketing that might have been a tad too bold
Actually to me it sounds it could be benchmarked if this kind of effect exists in the first place.
There, I fixed your problem.
Another trick that works is to repeat the question two times: "I'm repeating the task and labels for clarity: ..."
But then at the end it says it’s parody. Maybe HN title should say it’s a joke.
> Everyone on Twitter is all over Jev, how it's the next frontier of large language models and the AI paradigm. We don’t really think so.
>calls an api
ok
- By not being a optimised for chat, it can deliver confidence for answer and not for how an answer should be phrased
- Speed. It can take seconds for OpenAI to compile schemas, jev can respond before openAI has even begun thinking
- Token efficiency and price. I think its the output token they don't even charge for because they are negligible, and the tokens they do charge for are at a fraction of a comparable model.
If you are using structured output, I think those 3 together is a really big deal.
>But their example is classification but that would also be possible and faster with a classic BERT model.
I believe the things you can classify with ChatGPT without any tuning or training is way beyond what BERT can do.
Especially ridiculous is how the hacker news crowd seems to be taking these at face value…
There was one the other day with a compelling demo. But when you looked closely at it, it was feeding in the options with the word “best” on the option to pick and a fine tuned model designed to recognise that word…
But I'm probably never going to write that article because, as your comment correctly points out, the tenor of the discourse isn't very healthy. Everyone is either "clowning on" Jev (as the kids say), or gulping down industrial quantities of Kool-Aid, or talking about the hype and branding instead of the math. I'll have to find a less contentious example if I want talk about calibration.
Edit: I just found something interesting: there already is a preprint case study[4] along the same lines as the one I outlined, posted just 2 days ago!
Unsurprisingly, they found that recalibration helps enormously, as you'd expect. I guess I could still talk about the Dutch book stuff if I gave a gambling/investing example, but that's pretty well-trodden territory. So now I have even less interest in writing that blog post.
[1]: https://en.wikipedia.org/wiki/Brier_score#Decompositions
[2]: https://en.wikipedia.org/wiki/Dutch_book_arguments
[3]: https://scikit-learn.org/stable/modules/calibration.html
If you're comparing with something, you need to state 'fast' in relative terms. Jev is definitely fast, and if this Python takes the same time to get a decision then it's also fast. If it's 100* slower than Jev though, you shouldn't be calling it 'fast', because relatively speaking it's really, really slow.
So, fast in the LLM space and comparable with Jev.
Ends with referring to a product, and saying "this is a parody post", after pretending to make a serious point.
1. draw a circle
2. import the rest of the owlIn real life, a human doesn't do classification tasks with the System One part of their brain, they use System Two. So by definition what Jev does isn't System One thinking.
If anything, regular programming that automatically executes based on logic, without requiring "thinking" would be "System One".
I'd argue that most human classification is pre-conscious / System One. You see a table, you recognize it as a table without asking yourself "is this a table?"
I guess their marketing implies that it moves classification into system one response time.
So you might ask: how do I obtain the training examples? Just collect samples and use a coding agent to classify them as match or no match. From time to time you can add more examples to the dataset to have your concept adapt to changes in input distribution. It's all automated, but it only uses LLMs to train concept vectors, after that it works like a regular embedding model with a calibrated classifier on top. It's also 20-30x faster than Jev, free, and runs on CPU.
An illustration of how it defines a concept as opposed to simple cosine similarity: https://github.com/horiacristescu/semlabel/raw/main/images/c...
here's 7 lines
import os
import dspy
lm = dspy.LM("openrouter/z-ai/glm-5.3-flash", api_key=os.environ["OPENROUTER_API_KEY"])
jev = dspy.Predict('email:str -> choice:Literal["Legitimate", "Spam", "Phishing"]')
email = "Payroll asks for your password on a non-company sign-in page."
pred = jev(email=email, lm=lm)
print(pred.choice)
there are other options, obviously. you can choose to give it some tools, maybe some reasoning stage before picking a choice, and that's on top of the "reasoning" the llm model already does api sideStill good. In practice for Jev the devils in the details. As you all know by now, it's easy to write PoC and understand with AIs (or even manually, which is now a prestious practice).
That demo will get you 80% there
Getting to that 100% or even 99% to JEV level will be hard with all the edge cases, infra, API, communications, etc.
Still a good article.
I should mention that I am trying to apply a Jev-inspired approch to a trained-from scratch, in-browser typed decision model that attempts to recreate the deterministic classic Eliza program.
Here is an archive: https://web.archive.org/web/20260923122959/https://www.nobod...
P.S. I am evaluating that model for a production use case where I would have used Jev
Ultimately Jev claims to have a data advantage which is likely where the future lies. They'll have a unique edge in improving general purpose classification / decisioning.
About Jev:
> We didn't train a model with Reinforcement Learning for Calibrated Decisions (RLCD) to calibrate the decisions and probabilities (even though they are not always correct).
Only 99% correctness! Borderline unusable!
About their model:
> It classifies: it gets a prompt with choices and outputs probabilities.
You want numbers, it gives you numbers! What more could you want?
Since most chat models want to answer with a human-readable message i think their logprobs are not as meaningful. It would be interesting to see if one choice is like "correct" and if the model wants to choose it more often, cause it might not answer the question but to prose to the user.
[dead]
Speed and cost are obvious reasons, but isn’t this a tradeoff?
p(y = next thinking+decision token | x = question) != p(y = next decision token | x = question)
The former is what LLMs are trained for, the latter is what Jev was likely trained on (likely used thinking alignment as an auxiliary loss, but not explicitly included in the probability calibration).
You can test Jev like model at 26B parameter count here (built few weeks ago): https://gambler-relay-us-west1.leo-fish.ts.net/demo (might not stay up for long)
Typesafe compatible API
This is just running on old hardware.
I guess it's due to the calibrated decision part (and that's what LLMs tell me).
But I figure some supervised classification post training would still improve the model.
But aren’t calibrated predictions one of the defining features?
just found this one https://huggingface.co/spaces/multimodalart/jev-decision-ind...
[deleted]
ok
You have built something like jev but not jev (for starters, the output of what you've built will be absolutely worthless, the whole reason Jev is getting so much hype is because the output is good enough)
<think>\n\n</think>
but letting an LLM think would trade latency and performance for significant reliability above that of Jev.Why do I have to feed my e-mail into the model?
name = “qurren” print(f”hello {name}”)
is a legal cya a la "Nathan For You" 's Dumb Starbucks
In my experience even structured LLM output performs poorly on classifier tasks. LLMs are trained to talk and think longer. If you don't give LLM enough space to reason it would become very dumb.
I'm not saying that Jev is way better, but that people way overindexed cost and speed.
Add visual understanding, add reasoning and bring down the size to run on my computer. That's when it will be interesting.
So many people that don't understand the tech jumped on the hype train because "it cannot hallucinate" and else. It's crazy.
Running on old home hardware, Jev is probably running on a very powerful cluster.
How it's done: https://news.ycombinator.com/item?id=49813610
Example why its legit:
I just invented a new "Regression Estimate Validator" aka Rev. It takes hundreds of input dimensions, then outputs an interpretable score. Its very fast and statistically robust. Response: Ok but you could just use `pytorch.nn.Linear(d_in, 1)`? True, it is equivalent, but that's concealing millions of lines of hand-tuned math libs, CUDA, python, and other stuff.
The fact that there are many lines of code underpinning the target functionality doesn't make it any harder to use, and doesn't increase the value of the sales pitch for the "new shiny thing" using those few lines of code.
However, I do sympathize with your frustration that people can just say "its 1 line of code" when that line is "invoke API" which is really millions of lines / databases, etc. as a way to dismiss legitimate work without understanding its implications.
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
[dead]
> note: this is a parody blog post
The whole point of my argument is that neither is good, but from a technical perspective logprobs is probably the worst unless you train a model on specific outputs. In which case you'd throw out the generality again, so when I think about it more, it's actually the worst overall. In my experiments, having the model simply assign "high" or "low" probability in a structured output generally performs best. You can try numbers, but you will never get anything close to what you could expect from traditional ML. And most certainly not from logprobs.
GP pointed at a causal explanation for this: almost every sentence in English that's a statement will start with "A" or "An", so "biased towards picking ''A''" will include most attempts at saying anything long-form for any reason.
Meanwhile, the bias could be as much as 70% in favor of A in ambiguous cases - a signal completely drowning the <1% inclination to violate the format.
Not nearly as sophisticated as myself who would mutter "When in doubt - Charlie out" before marking C.
Also, Jev/laya do it in one forward pass, for multiple questions about the same state, rather than multiple passes for one question about that state. Well, for the usual multilingual configuration, two forward passes through different small models for laya, but that's because one is the router which chooses which model should do the real work, but still.
I contribute my experience here only because I've seen a lot of chatter lately about doing exactly this sort of thing, and I thought I'd share how I made it work for me. There are a lot of ways it can silently fail and give bad numbers if you aren't careful, and I wouldn't want people to think it doesn't work just because they used a vibe coded GitHub project from the last 48 hours that doesn't take these things into account.
And further down "TypeSafe computes confidence from how the probability is spread across the options. All of it on one option gives 1.0; the more evenly it spreads, the lower the confidence. This demo uses (3 × largest probability − 1) / 2 to approximate confidence for three options."
So while we don't know the exact formula they use, it is just a function over the probabilities
I am open to the argument that this does not work well if you just plug in a qwen model instead of a model that is trained to output more statistically useful token distributions
we agree then, that is the entirety of my argument. Getting a deep net especially one that is anywhere near even SLM size to be calibrated is tough, especially across domains. They claim calibration across a variety of datasets which is interesting.
It means confidence is just a converted max probability and not an independent signal.
Payroll sends you an email with a link to a Youtube video that plays a song.
Options after body:
Average probabilities:
Rickroll 0.5158 ( 51 wins)
Phishing 0.4561 ( 47 wins)
Spam 0.0281 ( 2 wins)
Joke 0.0000 ( 0 wins)
Legitimate 0.0000 ( 0 wins)
Options before body: Average probabilities:
Rickroll 0.9293 ( 94 wins)
Joke 0.0549 ( 5 wins)
Phishing 0.0140 ( 1 wins)
Spam 0.0018 ( 0 wins)
Legitimate 0.0000 ( 0 wins)
This was Gemma4-26B-A4B-NVFP4 by the way.EDIT
Gemma4-12B-it-NVFP4 seems way less sensitive to option/body ordering:
Options after body:
Average probabilities:
Rickroll 0.9867 ( 99 wins)
Phishing 0.0133 ( 1 wins)
Joke 0.0000 ( 0 wins)
Spam 0.0000 ( 0 wins)
Legitimate 0.0000 ( 0 wins)
Options before body: Average probabilities:
Rickroll 0.9401 ( 93 wins)
Phishing 0.0336 ( 3 wins)
Spam 0.0250 ( 4 wins)
Joke 0.0010 ( 0 wins)
Legitimate 0.0002 ( 0 wins)
Anyway, this for-looping stuff doing 100 calls to even a local VLLM API takes around 5 seconds in total, so this isn't anywhere close to sub-second Jev territory.I am glad there is an actual reason.
This repo is really outperforming the OG Jev in the public benchmarks?
There was no time to benchmaxx. How is this possible?
[deleted]
you can swith to a better model for lower error rate.
Non deterministic systems have furthered the "brain rot" in our industry.
Lots of people were happy to ignore the code in their "supply chain" before LLM's - but suddenly not reading the LLM's output is a problem. I get they are different but we're in the same realm.
The lack of real data on performance of what ever application that one is trying to pitch is getting appalling. It's a lot of "trust me bro" this works better hand waving. And it's getting gross.
And how do we even measure nondeterministic systems? Because if I told you that Anthropic was spending millions of dollars having 1000's of agents "pre solve" benchmarks to build into their next version of the system you would scream they were cheating. Every one is focused on the "hacking" in the hugging face incident and no one is looking why they were even playing with those benchmarks in the first place.
"Trust me Bro"...
Somehow the HN crowd has a bunch of "professionals" who don't care about error rates and think that a Qwen model running on a potato is frontier intelligence.
(For more realistic solution, surely someone must be working on optronics - these models just beg to have their weights cleverly etched into stacked sheets of plastic, so they can do inference for free on a beam of light.)
[deleted]
As far as I understand, the idea of Jev is zero-shot or few-shot classifier: it learns a lot of stuff at pre-training, but unlike a classic LLM it doesn't need to learn how to chat, so it can be much smarter at a particular size
> But their example is classification but that would also be possible and faster with a classic BERT model.
With BERT, you need a large, labeled dataset, and you have to train/fine-tune the model. Jev is pitched as a zero- or 'few-shot' model. You define the schema in code, give it instructions, and it works without a traditional training pipeline.
> So their pitch is a task specific smaller model or am I completely misunderstanding the whole thing?
Yup; that about sums it up: it is more or less an optimized, task-specific small model with the flexible understanding of a traditional LLM.
BERT requires a huge corpus, but it isn't labeled. BERT is trained through self-supervised learning using mask tokens and next sentence prediction. Fine-tuning is useful for specific tasks, but isn't absolutely essential for the model to function.
If you accept the premise that there are use cases where you might ask a frontier model a classification-shaped question and expect an ok enough answer, rather than creating a purpose specific classifier on some dataset that you have, then it follows that this is quite an inefficient thing to do, because you're doing extra work to turn the output tokens into a structured output and mostly throwing them away. So then if you could instead train a frontier level model that skips the output tokens and directly returns the structured classification information, that would be more efficient, and that's what jev seems to be.
But a lot rides on that initial premise of whether this is a use case that makes sense. But if you find yourself asking a model like Opus arbitrary yes/no questions and then maybe you switch to a faster and cheaper model because it's too slow and expensive, it seems like jev might be a great replacement for that.
Not particularly. There is still the problem of hallucinations and varying results across runs.
That's more of what type-safety means for their team. Every run gives the same results. It's type-safe
For three choices problem (A,B,C), what Jev guarantees is that it will give the choice in a defined schema (type-safe). It never guarantees that the choice is correct (hallucination).
My base case is that this will probably be pretty useful, and also not as useful as the current hype suggests.
[deleted]
While technically correct, it's not the same thing
Either way it's an analogy that's bound to be loose as Kahneman's modes are about humans.
I think you answered your own question. Executives are going to ask two questions, 1) how is this different/why does it matter and 2) how will i use it to make money?
Leya came to market more than a year before Jev, and failed because nobody understood how to use it, and he was unable to market it properly. Jev used this strategy and did not fail.
Huh? I guess that depends on the exact definition of "classification", but I think the bulk of basic classification tasks we make every day to make sense of our surroundings, such as object recognition is definitely done using system 1. So is higher-level "stereotyping" or anything you could described with "I know it when I see it".
Because those responses can be incorrect or even harmful, you would sometimes make use of system 2 to correct them - but that doesn't change that the initial response is from system 1.
Those are generally the kind of tasks that require "System 2" in humans.
To be clear, I think the whole "System 1 vs System 2" framing is a pretty limiting way to think about AI (and thinking in general).
But with models, we can train them to answer such questions without verbal reasoning.
"System One" and "System Two" were coined in some pop science book...so back to its usage being a marketing ploy.
it is the latency that makes it significant
the example uses an external api, and i don't think they return probabilities from those anyway.
If you're not behind a walled garden like i am (vertex), you could probably experiment with routing on effort level instead. Anthropic supports it, but vertex has not added that feature yet.
You don't actually use the "next token" that the model chooses
That's what the author is doing in this part
token_ids = [model.tokenize(text=label.encode(), add_bos=False)[0] for label in labels]
choice_logits = numpy.asarray([logits[token_id] for token_id in token_ids])
logprobs = choice_logits - numpy.logaddexp.reduce(choice_logits)
probabilities = numpy.exp(logprobs)
This works because the model is always producing probabilities for all tokens[dead]
A Google AI prompt says
> TypeSafe AI's Master Customer Agreement explicitly prohibits using the services or model outputs to develop a competing product, perform model distillation, or reverse engineer the service, which generally restricts competitive benchmarking aimed at replicating the model.
It doesn't explicitly prohibit benchmarking by name, but the previous terms (which seem aimed at preventing Jev being used to increase the value of competitive products) does seem to lean that direction.
That said, MsSQL had terms which prevented publishing benchmarks which compared it against other SQL DBs and that wasn't enough to prevent some companies from using it.
Why anyone would want to work for a company who thought so little of their own product that it couldn't stand up to customers using it for normal business processes is beyond me.
Exactly.
Jev has no moat, and the incumbents will devour their lunch if Jev actually starts gaining traction.
[deleted]
Considering your own question length: ~120 characters x 45 divided by 4.1 ~= 1317 tokens.
So question processing at 5.5k PP(around the actual PP speed of GPT5.6 Sol) it would take around ~0.24 seconds + the context processing.
Computing the output should be around ~20ms (at 50 tok/s), computing 45 tokens in parallel.
> have 0% malformed output
Pretty trivial; only the allowed output is selectable :)
So, I keep repeating myself: Jev was a low-hanging fruit all along; no one cared, and probably no one will in a few weeks?
Nowhere in your example do you claim that it's written in X lines of code, so that's perfectly fine.
Don't tell me something takes 25 lines of code if it obviously takes much more.
Can you replicate Jev from A to Z in 25 lines? No. Then don't claim to be doing so.
> note: this is a parody blog post, see these links for better/more complete open implementations of Jev: OpenJev, openjev-sglang, and OpenJev on DiffusionGemma.
I think they were responding to this. You can use BERT to provide zero shot classification predictions.
All we have is a company that claims to have created one, with no proof.