[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"article-detail":3},{"article":4,"alternate":35,"related":40,"latest":49},{"orderNumber":5,"showImage":6,"title":7,"metaDescription":8,"jumpUrl":9,"content":10,"metaKeywords":11,"withAllowSearch":12,"modified":13,"viewCount":14,"id":15,"lang":16,"slug":17,"dpTemplateId":18,"thumbnail":6,"withHot":19,"withLeadNews":19,"author":20,"created":13,"withTop":19,"highlightContent":21,"userId":22,"highlightTitle":7,"commentStatus":12,"commentCount":5,"thumbnailToContent":19,"withRecommend":19,"metaTitle":23,"editMode":24,"siteId":5,"user":25,"authorEn":20,"status":33,"categoryId":34,"summary":28},0,"\u002Fattachment\u002F20260929\u002Fdd675237c2fe42d6b8d4823bbdd32e30.webp","Jev AI Model: Judgment Without Words","Jev is an AI judgment model that returns probabilities instead of text. Here is how it works, what tests found, and why the paradigm-shift crown is premature.","https:\u002F\u002Fpoly-ai.chat\u002Fmediasync-claw","\u003Cp style=\"margin:0 0 18px;\">Tech hype runs in eerily familiar cycles. In September 2026, the whole of Silicon Valley and the open-source community went wild over a model named Jev — a self-proclaimed \"System One\" model promising something genuinely different: direct judgment, fast.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">For the past three or four years, we have lived with large language models that behave like rambling essayists. Before answering a simple \"yes or no,\" they generate hundreds of thinking tokens, then slowly produce a JSON conclusion. Jev skips that entirely: it promises probabilities, not paragraphs — and it is fast enough to feel instant.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Developers building AI agents went crazy for it. When your agent needs to ask \"is this memory relevant?\", \"which tool should handle this request?\", or \"does this incident need a human review?\", a general LLM burns tokens and latency on every one of those small questions. Jev cuts into a very real pain point.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">But is it truly disruptive? Strip away the packaging, and the real question is: how much of the judgment inside the black box comes from the underlying language model, and how much comes from TypeSafe's supposedly special training? To answer that, we need to peel Jev apart layer by layer.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cimg src=\"\u002Fattachment\u002F20260929\u002Ff0fa37222d4e4101ab992bb69d805e1d.webp\" alt=\"Jev AI model versus token-heavy LLMs — text generation versus direct judgment\">\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Judgment models are older than you think\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The direction Jev represents has a long history. Back in 2018, Google released BERT — an encoder-only architecture that reads text bidirectionally and was trained to fill in blanks. Although decoder-only models like ChatGPT later became mainstream, BERT kept a real advantage on classification tasks: seeing the whole context lets it build complete feature representations, topped with a simple classification layer.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Modern LLMs have often been quietly trained as classifiers too. In 2019, OpenAI's \"Fine-Tuning Language Models from Human Preferences\" began teaching models human preferences. By 2020, work on text summarization made the pipeline explicit: human annotators compare two summaries, a reward model learns which one humans prefer, and human judgment is compressed into a scoring function — the model reads question and answer, and its output layer directly emits a score. No essay required.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">That is RLHF — the very mode of learning Jev rails against. \"Inherit language understanding but do not generate text\" is not Jev's original idea. The 2023 paper \"Let's Verify Step by Step\" pushed supervision deeper into each step, and by 2025–2026 the field had Galileo's Luna-2 (a small model trained as a single-token classifier) and Skywork-Reward-V2's reward-model series. From an algorithmic standpoint, making an LLM emit candidate scores from its hidden representations instead of words is not hard — there are at least three ways to do it, as the replication projects later showed.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">So what does Jev actually add? Two things, according to its architecture notes and official claims.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">First, a more general interface, thanks to a new post-training method. Past scorers were trained per task. Jev tries to be a universal probability interface — and its probabilities target \u003Ci>facts\u003C\u002Fi>, not human preferences. TypeSafe calls the method RLCD (reinforcement learning for calibration-based decision-making), putting probability calibration at the center of training. Traditional RLHF models human preference probability; Jev claims to model real-world event probability.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Second, ruthless parallelization. Jev changes how information flows through the model. It can answer up to 250 questions against the same source material, as long as the questions have no sequential dependencies — massive, batched parallelism.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Reconstructing Jev from tests and replications\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">TypeSafe has published the interface but not the computation graph. The interface accepts a State (the shared material — say, a long customer complaint or system log) plus multiple Questions. To keep things standard, questions are squeezed into three primitives:\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Noul: a yes\u002Fno question returning a probability between 0 and 1 (is this urgent? → 0.95).\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Choice: pick one among candidates, returning a probability distribution (tech support 0.8, billing 0.2).\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Score: rate a degree, returning probabilities per level plus a weighted score (customer anger 4.5\u002F5).\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Jev claims to read the State only once; all questions are then judged in parallel within a single request. Questions cannot peek at each other — if question two depends on question one, you must send two requests. TypeSafe encourages \"speculative fan-out\": ask everything at once, discard what the downstream code doesn't need.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cimg src=\"\u002Fattachment\u002F20260929\u002Faa8fc512360b47ecb3091054db7ee3f3.webp\" alt=\"Jev primitives — yes-no probability, choice distribution, score rating\">\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Black-box testing by Archer Hume backs the \"read once\" claim. Billing showed 268 input tokens for a single yes\u002Fno question versus 276 for two — the increase was only the added question text, not the shared State. Response time stayed near-flat until nearly a hundred questions. The open-source replication Kev implements exactly this: process the State once, freeze the intermediate result in the KV cache, and let all questions share it.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Are the parallel questions truly isolated? Hume ran a clever \"code-word experiment\": he hid \"the code word is ZEBRA-7741\" inside question A, then asked question B to identify a code word mentioned in another question. Jev returned the correct code word with probability 0.00. Move the code word into the shared State, and question B jumped to 0.90+. That is strong evidence of physical isolation between questions — shared material is visible to all, but neighboring questions cannot peek at each other.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">On implementation, Kev uses attention masking (zeroing out other questions' regions while one question computes) and, for base models whose attention masks can't separate branches, independent branch reuse — the model reads the State, then forks into parallel \"highways,\" each inheriting the frozen shared memory.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">How are options scored inside a single choice question? The most traditional method is a linear head plus softmax — the \"black-room blind review\": each candidate is scored in absolute isolation by a fixed rubric, then softmax converts the raw scores into percentages. Under that model, adding a nonsense distractor like \"bad weather\" could never change the relative odds of \"finance\" versus \"tech.\" Hume's tests proved it does — adding a distractor shifted the relative probabilities. That kills the blind-review hypothesis.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">So candidates must \"see\" each other before the final score. Open-source projects offer two blueprints: Kev's pointer head (a group interview where candidates are compared directly) and NanoJev's inter-candidate attention module (each candidate's features are first encoded separately, then a small attention module compares them together — the distractor joins the meeting and reshuffles the weights).\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Why would Jev go to such lengths to make options compete in the underlying code? Because in real business, the right answer is often relational: options are hidden clues. Ask \"where is the Eiffel Tower?\" with candidates A. Europe, B. France, C. Paris — seeing all three at once reveals that the test is about the highest precision of location, not rough geography.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Finally, the result: Jev returns probability numbers directly, without autoregressive text generation. External probing confirms this — when Hume inflated the candidate list from two to two hundred, the response text grew long but processing time did not scale. The extraction method varies: openjev reads the raw logits of candidate tokens at the answer position; Kev's pointer head emits comparison scores directly; minojev relies on a shared scoring module. Any of them lets the model skip the slowest step — predicting word by word — and pull an exact probability out at the end.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The architecture, in the end, is not complicated. There is little in it that deserves the word \"paradigm shift.\" It is a well-engineered optimization for a specific scenario — and only that.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Accuracy comes from post-training\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Speed is architecture; Jev's accuracy is claimed to come from RLCD post-training, which is a closed box. The open-source community has been guessing at its recipes.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Synthetic data is the simplest route: the Hmm replication had DeepSeek V4.1 generate hundreds of work scenarios (refunds, troubleshooting, retrieval relevance, email routing), each with materials, questions, candidates, judgment standards, and answers — with quality gating by re-answering hidden answers three times. Kev took existing datasets (news classification, sentiment, textual entailment) and converted them into \"material + question + candidate\" format. Kev-4B's September 24 version went further, building questions from 5,219 real consumer-finance complaints — keeping a label only when two different teacher models agreed with the original filing.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The replications also craft paired trap questions: two questions with identical rules but one key name swapped (a signer with authority, Mira, versus one without, Noah) flips the answer. This stops reward hacking and forces the model to learn deep representations of the question–answer link.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Is LoRA plus distillation enough? Most replications use LoRA fine-tuning plus teacher distillation to raise the probability of the correct option. Winnow, for example, LoRA-tunes a Gemma 4 12B instruction model with two kinds of supervision: the standard answer (push up the correct option's probability) and a teacher's full probability distribution (distill the distribution). Both losses use cross-entropy. LoRA, distillation, and cross-entropy are generally adequate for a probability-prediction task — but distillation learns the teacher's probabilities, not the real-world probabilities RLCD claims to target. How to bridge that gap is an open question for the replications.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Calibration may be the actual trump card. Temperature calibration — tuning overall confidence down so probabilities look humble — does not change option ranking. After calibration, Kev-9B's calibration error dropped from about 10.6 points to 4.2 points, with accuracy unchanged. It is a blunt instrument (it lowers global confidence rather than distinguishing when the model \u003Ci>should\u003C\u002Fi> be confident), but if Jev achieves better calibration, it may indeed have a genuine trick.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cimg src=\"\u002Fattachment\u002F20260929\u002F3c68669a2f04428fa66eac3393f5ca7a.webp\" alt=\"Jev parallel fan-out — one shared state answering many independent questions\">\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>Where Jev actually fits\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">Jev has real use. It returns fast decisions to the tasks that only ever needed fast decisions: request classification and routing, retrieval-result ranking, and the many checklists inside agent flows — customer-service triage, product categorization, feedback analysis, data labeling.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The benchmarks show some generality: Nimble's team tested it across 13 groups of 3,880 public samples (fact-checking, intent routing, textual entailment, content moderation, medical QA) and Jev averaged 76.0% accuracy. But true generality has two gates.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The first is hard tasks. Judging \"does this refund request exist\" is surface semantics; judging \"should this refund be approved\" means checking dates, computing limits, and comparing clause priorities — multi-step dependency. On JevBench's hard questions, Jev's weakness shows. Against GPT 5.6 Luna over 616 valid question pairs, the knowledge gap was only 1.6 points, but math-and-reasoning widened to 19.3 points and code to 20 points.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The second gate is generalization. In identical tasks, Jev's judgment is easily disturbed by how information is presented: OpenProse moved the key relations from the front of a text to the middle — same facts, same question — and accuracy halved from 80.5% to 40.9%. On JevBench v1.4's 308 closed hard questions, Jev collapsed from 86.6% on public questions to 36.7%, while thinking-mode DeepSeek V4.1 Flash held 94.8%. Transfer to real business is mixed: on Agent Journal's prompt-injection detection, Jev improved from 83.65% to 95.58% on new external data — but Scarif Labs found Jev's AUROC for judging software-update safety dropped from 0.851 to 0.605 (near random) when moving ecosystems.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">The honest position: use Jev where standards are clear, evidence is concentrated, and the judgment can be checked or corrected — early triage. Keep math and date comparisons in code, as TypeSafe itself advises, and minimize multi-layer dependencies. And watch the cost ledger: in one GitHub agent memory-retrieval experiment, adding Jev as a relevance judge pushed total latency from 649 ms to 1087 ms and more than doubled cost per thousand calls — if the main model still gets called anyway, the added judge must save enough downstream work to pay for itself.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">\u003Cstrong>The crown question\u003C\u002Fstrong>\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">What is Jev actually learning? A judgment model's representation should support effective judgment from new facts — among the hardest representations to learn. \"I want a refund\" is language understanding; \"should this refund be approved under this policy\" requires mapping facts to clauses and handling exceptions — mostly the underlying language model's ability. But \"will refunding keep this customer?\" requires predicting consequences of actions: which facts matter, under what conditions, and whether those relationships survive a change of scene. That is the part that needs training — and, for humans, comprehensive judgment is one of the hardest things to learn.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">So we can fairly call Jev a clever engineering tool. In business pipelines with clear rules and complete materials, it cuts latency and compute costs — real value. But until it proves it has learned a general rule of judgment, putting the \"paradigm shift\" crown on it is premature.\u003C\u002Fp>\n\u003Cdiv class=\"dp-template-card\" style=\"border-radius:8px;box-shadow:0 2px 8px rgba(0,0,0,0.1);margin:10px 0;max-width:100%;overflow:hidden;width:100%;\">\n \u003Ca style=\"display:block;text-decoration:none;\" href=\"https:\u002F\u002Fpoly-ai.chat\u002Fmediasync-claw\" target=\"_blank\">\u003Cimg class=\"image_resized\" style=\"display:block;height:auto;max-width:100%;width:100%;\" src=\"\u002Fattachment\u002F20260824\u002Fdb1f58e4d6e24f2f8dcccc811df6e1a8.png\" alt=\"db1f58e4d6e24f2f8dcccc811df6e1a8\">\n  \u003Cbutton style=\"align-items:center;background-color:#784fe2;border-radius:0 0 4px 4px;border-style:none;color:#ffffff;cursor:pointer;display:flex;font-family:Times New Roman;font-size:20px;height:40px;justify-content:center;padding:0;width:100%;\">Experience Now\u003C\u002Fbutton>\u003C\u002Fa>\n\u003C\u002Fdiv>\n\u003Cp style=\"margin:0 0 18px;\">References:\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">[1] TypeSafe Jev official interface documentation and System One announcement, 2026.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">[2] \"Fine-Tuning Language Models from Human Preferences,\" OpenAI, 2019.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">[3] \"Let's Verify Step by Step,\" 2023.\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">[4] SUNSHINE-equivalent benchmark sources cited in the original analysis (JevBench, Nimble, Agent Journal, Scarif Labs, OpenProse).\u003C\u002Fp>\n\u003Cp style=\"margin:0 0 18px;\">[5] All illustrations are AI-generated.\u003C\u002Fp>","yunpoly, aipollo, mediasync-claw, Jev, System One model, AI judgment model",true,"2026-09-29 00:10:35",75,12601,"zh_CN","jev-judgment-model-zvrn",23,false,"yunpoly","Tech hype runs in eerily familiar cycles. In September 2026, the whole of Silicon Valley and the ope...",18,"Jev AI Model: Judgment Without Words | yunpoly","html",{"emailStatusOk":19,"statusLocked":19,"mobileStatusOk":19,"englishNickname":26,"nickname":26,"statusReg":19,"id":22,"created":27,"sourceString":28,"avatar":29,"url":30,"statusOk":19,"detailUrl":31,"username":32},"Science Guide Wwai","2025-12-29 12:22:32","","\u002Fstatic\u002Fcommons\u002Fimg\u002Favatar.png","\u002Fuser\u002F18","\u002Fadmin\u002Fuser\u002Fdetail\u002F18","panhb","normal",72,{"id":36,"slug":37,"title":38,"lang":39},12637,"c69497-lactose-intolerance-gene-benefit-diabetes-risk","乳糖不耐竟是「基因紅利」？研究揭祕：拉肚子竟能防糖尿病","zh",{"prev":41,"next":45},{"id":42,"slug":43,"title":44,"categoryId":34},12604,"e21b09-viral-video-child-confronts-mother-grandfather","The Mirror of Childhood: How One Boy’s Brutal Honesty Exposed Parental Hypocrisy",{"id":46,"slug":47,"title":48,"categoryId":34},12599,"deepseek-thinking-upgrade-nxqt","DeepSeek's Thinking Upgrade: One Model, Two Minds",[50,55,60,65],{"id":51,"slug":52,"title":53,"thumbnail":54,"categoryId":34},12664,"2f0b3c-ai-snake-romance-gender-debate","When the Fiancée Is a Snake: AI Anime Sparks Debate on Gender Roles and Absurd Romance","https:\u002F\u002Fcdn.banyunjuhe.com\u002Fattachment\u002F20260930\u002Fabd307865c604c1fa501ee8fe6cfff73.png",{"id":56,"slug":57,"title":58,"thumbnail":59,"categoryId":34},12662,"b86dc5-ai-scheming-doubao-debate","AI Auctions Spark Debate: Is 'Doubao' the Ultimate Schemer?","https:\u002F\u002Fcdn.banyunjuhe.com\u002Fattachment\u002F20260930\u002Fcb0922ad05854444856a5e6ff36be9b5.png",{"id":61,"slug":62,"title":63,"thumbnail":64,"categoryId":34},12660,"b5e00a-ai-short-drama-tutorial-controversy","The Rise of the AI \"Evil Cultivator\": A New Era of Digital Content Creation","https:\u002F\u002Fcdn.banyunjuhe.com\u002Fattachment\u002F20260930\u002Fdb61bf4e6fd74f36844203afb950ce59.png",{"id":66,"slug":67,"title":68,"thumbnail":69,"categoryId":34},12658,"0cdfe4-ai-supercar-design-apollo-bugatti-ferrari","Can AI-Designed Supercars Outshine Legacy Brands?","https:\u002F\u002Fcdn.banyunjuhe.com\u002Fattachment\u002F20260930\u002F67262dae9f7048cb8ad5d44f5a0ce3e5.png"]