LA Evaluation & ObservabilityMIT
Latitude includes an optional Jev preclassifier for conversation checks and their selection records.
Judges which checks apply and can add checks when thresholds and rate limits permit.
Records models, thresholds, latency and selection reasons alongside the baseline.
Evaluation & Benchmarks Voice & Conversation
NI Evaluation & ObservabilityMIT
A local MCP code-quality reviewer returning structured scores to coding Agents.
Jev scores correctness, complexity, tests and security; code ranks areas to improve.
Compares scores across checkpoints while leaving code changes to the primary Agent.
Coding Agents MCP & Integrations Evaluation & Benchmarks
LD Evaluation & ObservabilityMIT
An optional Jev judgment module in Taskuary for checking user-defined conditions on task state.
Turns conditions into yes/no probability questions and returns local-threshold verdicts plus probabilities.
Adds structured checks to task outcomes without making Jev the controller of the whole messaging system.
Evaluation & Benchmarks Typed Decisions
SU Evaluation & ObservabilityMIT
A quality and coverage CLI for coding agents: Jev assesses source properties while local coverage highlights testing targets.
Asks Jev about file-quality properties; code composes scores and ordering.
Breaks scores into named properties and caches results by content.
Coding Agents CLI & Git Gates Evaluation & Benchmarks
AL Evaluation & ObservabilityMIT
A film-and-TV attribute-scoring experiment inside GoodWatch comparing Jev question designs and batch sizes.
Asks whether predefined traits are present or how strongly they appear, recording scores, latency and Token usage.
Compares rating scales, input variants and batching over a frozen sample.
Creative & Multimedia Evaluation & Benchmarks
IA Evaluation & ObservabilityMIT
Compares Jev with other structured-output models on shared application tasks, recording errors, latency, Tokens, and estimated cost.
Maps the same tasks to Jev Choice/Noul questions and normalizes answers to a shared result format.
Preserves comparison methods and results for inspecting model differences.
Evaluation & Benchmarks Typed Decisions
QK Evaluation & ObservabilityMIT
Keeps an execution ledger for Claude Code and Codex CLI to check for passing validation after edits.
Jev can identify completion claims and semantic-rule issues; stop gates depend on ledger facts and local rules.
Separates execution evidence from model opinion rather than letting Jev alone certify completion.
Coding Agents CLI & Git Gates Evaluation & Benchmarks
MI Evaluation & ObservabilityNo license specified
A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.
Experiments send action candidates or typed questions to Jev, then execute or record the answers.
Includes source, experiment notes and some offline replays for comparing decision designs.
Evaluation & Benchmarks Games & Simulation Browser Automation
AB Evaluation & ObservabilityApache-2.0
A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.
Runs the same labeled text tasks through both backends and records probabilities, latency and failures.
Helps examine task-specific accuracy and whether confidence scores support chosen thresholds.
Evaluation & Benchmarks Classification & Ranking
KA Evaluation & ObservabilityNo license specified
Interactive Jev experiments for support-routing previews and 3D driving simulations.
Judges support messages or chooses lanes and target speed from structured simulated sensors.
Shows inputs, probabilities and resulting behavior together.
Games & Simulation Classification & Ranking Evaluation & Benchmarks
AB Evaluation & ObservabilityMIT
A toolkit for evaluating Jev probabilities on labeled data, selecting confidence thresholds and checking model drift.
Runs fixed questions and measures accuracy, calibration, coverage and escalation rates.
Connects threshold selection and model-change checks to reports and CI.
Evaluation & Benchmarks Typed Decisions
AN Evaluation & ObservabilityMIT
Compares Jev, dedicated rerankers and chat models on the same retrieved passages.
Ranks candidate passages with Choice, Noul and rubric scores, then computes retrieval metrics.
Publishes raw responses, scoring code and per-dataset results for inspection.
Evaluation & Benchmarks Search & Retrieval Classification & Ranking
WO Evaluation & ObservabilityMIT
Benchmarks Jev on chess moves and identifying which game NPC a player addresses.
Selects legal chess moves or judges whether an utterance addresses each NPC.
Publishes labeled data, raw requests and responses, and evaluation code.
Evaluation & Benchmarks Games & Simulation
Y0 Evaluation & ObservabilityMIT
A word-level generation experiment that asks Jev to select words or verify locally drafted continuations.
Choice selects the next word; Noul evaluates candidate chunks and stopping conditions.
Exposes the limitations of using a decision model as a text generator.
Evaluation & Benchmarks Typed Decisions
DO Evaluation & ObservabilityMIT
Turns AGENTS.md preferences into rules checked by Jev against hunks, staged files or PRs.
Jev classifies change evidence against configured rules; code maps answers to review outcomes.
Feeds semantic-rule findings to coding Agents without replacing type checks, tests or security audits.
Coding Agents CLI & Git Gates Evaluation & Benchmarks
AD Evaluation & ObservabilityMIT
A research chat decoder that repeatedly asks Jev to choose words or phrases and assembles them in code.
Compares stepwise Choice decoding with selection from complete candidate replies.
Provides decoder methods, experiment traces and documented failure cases.
Evaluation & Benchmarks Typed Decisions
RI Evaluation & ObservabilityMIT
An independent Jev 1.13.0 behavior study recording successes and failures across question framing, input conditions and games.
Sends controlled variants of fixed tasks and records choices, probabilities and raw request-response evidence.
Lets readers inspect individual cases rather than infer broad capability from simple-task success.
Evaluation & Benchmarks Games & Simulation
SA Evaluation & ObservabilityNo license specified
A research repository tracking Jev claims and limitations, with calibration experiments and runnable examples.
Calls Jev on defined questions and labeled cases, then analyzes errors, calibration, and difficulty effects.
Links research claims to experiment code, data, and an evidence ledger.
Evaluation & Benchmarks Typed Decisions
NA Evaluation & ObservabilityNo license specified
Frontend QA that uses Jev to choose browser actions and checks contracts through DOM, HTTP and database evidence.
Jev selects observed controls and operations; test code owns expected values and pass criteria.
Records exploratory behavior separately from contract acceptance.
Browser Automation Evaluation & Benchmarks
TO Evaluation & ObservabilityApache-2.0
A benchmark using Jev to attribute multi-Agent failures to an Agent, step and error type.
Builds candidate sets from traces and submits three choice questions.
Provides evaluation scripts and author results; some baselines generate answers while Jev selects candidates.
Evaluation & Benchmarks Classification & Ranking
4E Evaluation & ObservabilityNo license specified
Compares Jev and OpenRouter models on labeled tasks for accuracy, calibration, latency and cost.
Collects judgments on matched tasks and computes confidence intervals and repeat-input stability.
Publishes data processing, runner and statistics code with model-specific results.
Evaluation & Benchmarks Classification & Ranking
OM Evaluation & ObservabilityMIT
A Windows PowerShell tool for auditing recorded Codex execution evidence with :jev.
Sends selected records to Jev for judgments about execution claims and evidence sufficiency.
Only reads and sends records on explicit invocation; judgments do not replace real tests.
Coding Agents CLI & Git Gates Evaluation & Benchmarks
PI Evaluation & ObservabilityNo license specified
A Jev title-and-abstract screening experiment compared with Cohen Abstract Triage labels for an ADHD review.
Asks Choice and Noul questions about eligibility, then combines them into include or exclude decisions.
Records metrics for specific dataset slices and question designs to examine screening errors.
Domain Workflows Search & Retrieval Evaluation & Benchmarks
PO Evaluation & ObservabilityMIT
A Hermes goal-checking skill that saves requirements, creates a run prompt and checks completion evidence.
Local checks handle machine-verifiable requirements; Jev assesses the remaining semantic conditions.
Connects goals with checkable criteria without replacing actual acceptance evidence with model opinions.
Evaluation & Benchmarks MCP & Integrations