Jev Agent
awesomejevRADAR
System-1 Radar Ready487 repos indexed

Every open-sourceJev project, in one place.

A source-verified directory of 487 Jev projects — browser agents, model routers, SDKs and MCP servers — each one showing the exact decision it hands to Jev. Curated by the Awesome Jev community, mirrored here under MIT.

Turn state into choices.

CHOICESCORENOUL
487Verified projects
17Architecture categories
20Topics
Updated09/22, 13:05
Popular projects
By current stars
Popular searches

Evaluation & Observability29

29 matching projects
Evaluation & ObservabilityMIT

Latitude includes an optional Jev preclassifier for conversation checks and their selection records.

WHERE JEV FITS

Judges which checks apply and can add checks when thresholds and rate limits permit.

Records models, thresholds, latency and selection reasons alongside the baseline.

4,665392

GitHub
Evaluation & ObservabilityMIT

A local MCP code-quality reviewer returning structured scores to coding Agents.

WHERE JEV FITS

Jev scores correctness, complexity, tests and security; code ranks areas to improve.

Compares scores across checkpoints while leaving code changes to the primary Agent.

19619

GitHub
Evaluation & ObservabilityMIT

An optional Jev judgment module in Taskuary for checking user-defined conditions on task state.

WHERE JEV FITS

Turns conditions into yes/no probability questions and returns local-threshold verdicts plus probabilities.

Adds structured checks to task outcomes without making Jev the controller of the whole messaging system.

11622

GitHub
Evaluation & ObservabilityMIT

A quality and coverage CLI for coding agents: Jev assesses source properties while local coverage highlights testing targets.

WHERE JEV FITS

Asks Jev about file-quality properties; code composes scores and ordering.

Breaks scores into named properties and caches results by content.

934

GitHub
Evaluation & ObservabilityMIT

A film-and-TV attribute-scoring experiment inside GoodWatch comparing Jev question designs and batch sizes.

WHERE JEV FITS

Asks whether predefined traits are present or how strongly they appear, recording scores, latency and Token usage.

Compares rating scales, input variants and batching over a frozen sample.

382

GitHub
Evaluation & ObservabilityMIT

Compares Jev with other structured-output models on shared application tasks, recording errors, latency, Tokens, and estimated cost.

WHERE JEV FITS

Maps the same tasks to Jev Choice/Noul questions and normalizes answers to a shared result format.

Preserves comparison methods and results for inspecting model differences.

375

GitHub
Evaluation & ObservabilityMIT

Keeps an execution ledger for Claude Code and Codex CLI to check for passing validation after edits.

WHERE JEV FITS

Jev can identify completion claims and semantic-rule issues; stop gates depend on ledger facts and local rules.

Separates execution evidence from model opinion rather than letting Jev alone certify completion.

292

GitHub
Evaluation & ObservabilityNo license specified

A MoonBit and TypeScript Jev playground covering games, browsers, command risk and small languages.

WHERE JEV FITS

Experiments send action candidates or typed questions to Jev, then execute or record the answers.

Includes source, experiment notes and some offline replays for comparing decision designs.

190

GitHub
Evaluation & ObservabilityApache-2.0

A benchmark comparing Jev and GLiNER on text classification, probability calibration and selective automation.

WHERE JEV FITS

Runs the same labeled text tasks through both backends and records probabilities, latency and failures.

Helps examine task-specific accuracy and whether confidence scores support chosen thresholds.

163

GitHub
Evaluation & ObservabilityNo license specified

Interactive Jev experiments for support-routing previews and 3D driving simulations.

WHERE JEV FITS

Judges support messages or chooses lanes and target speed from structured simulated sensors.

Shows inputs, probabilities and resulting behavior together.

114

GitHub
Evaluation & ObservabilityMIT

A toolkit for evaluating Jev probabilities on labeled data, selecting confidence thresholds and checking model drift.

WHERE JEV FITS

Runs fixed questions and measures accuracy, calibration, coverage and escalation rates.

Connects threshold selection and model-change checks to reports and CI.

100

GitHub
Evaluation & ObservabilityMIT

Compares Jev, dedicated rerankers and chat models on the same retrieved passages.

WHERE JEV FITS

Ranks candidate passages with Choice, Noul and rubric scores, then computes retrieval metrics.

Publishes raw responses, scoring code and per-dataset results for inspection.

50

GitHub
Evaluation & ObservabilityMIT

Benchmarks Jev on chess moves and identifying which game NPC a player addresses.

WHERE JEV FITS

Selects legal chess moves or judges whether an utterance addresses each NPC.

Publishes labeled data, raw requests and responses, and evaluation code.

51

GitHub
Evaluation & ObservabilityMIT

A word-level generation experiment that asks Jev to select words or verify locally drafted continuations.

WHERE JEV FITS

Choice selects the next word; Noul evaluates candidate chunks and stopping conditions.

Exposes the limitations of using a decision model as a text generator.

51

GitHub
Evaluation & ObservabilityMIT

Turns AGENTS.md preferences into rules checked by Jev against hunks, staged files or PRs.

WHERE JEV FITS

Jev classifies change evidence against configured rules; code maps answers to review outcomes.

Feeds semantic-rule findings to coding Agents without replacing type checks, tests or security audits.

40

GitHub
Evaluation & ObservabilityMIT

A research chat decoder that repeatedly asks Jev to choose words or phrases and assembles them in code.

WHERE JEV FITS

Compares stepwise Choice decoding with selection from complete candidate replies.

Provides decoder methods, experiment traces and documented failure cases.

30

GitHub
Evaluation & ObservabilityMIT

An independent Jev 1.13.0 behavior study recording successes and failures across question framing, input conditions and games.

WHERE JEV FITS

Sends controlled variants of fixed tasks and records choices, probabilities and raw request-response evidence.

Lets readers inspect individual cases rather than infer broad capability from simple-task success.

30

GitHub
Evaluation & ObservabilityNo license specified

A research repository tracking Jev claims and limitations, with calibration experiments and runnable examples.

WHERE JEV FITS

Calls Jev on defined questions and labeled cases, then analyzes errors, calibration, and difficulty effects.

Links research claims to experiment code, data, and an evidence ledger.

30

GitHub
Evaluation & ObservabilityNo license specified

Frontend QA that uses Jev to choose browser actions and checks contracts through DOM, HTTP and database evidence.

WHERE JEV FITS

Jev selects observed controls and operations; test code owns expected values and pass criteria.

Records exploratory behavior separately from contract acceptance.

20

GitHub
Evaluation & ObservabilityApache-2.0

A benchmark using Jev to attribute multi-Agent failures to an Agent, step and error type.

WHERE JEV FITS

Builds candidate sets from traces and submits three choice questions.

Provides evaluation scripts and author results; some baselines generate answers while Jev selects candidates.

20

GitHub
Evaluation & ObservabilityNo license specified

Compares Jev and OpenRouter models on labeled tasks for accuracy, calibration, latency and cost.

WHERE JEV FITS

Collects judgments on matched tasks and computes confidence intervals and repeat-input stability.

Publishes data processing, runner and statistics code with model-specific results.

10

GitHub
Evaluation & ObservabilityMIT

A Windows PowerShell tool for auditing recorded Codex execution evidence with :jev.

WHERE JEV FITS

Sends selected records to Jev for judgments about execution claims and evidence sufficiency.

Only reads and sends records on explicit invocation; judgments do not replace real tests.

10

GitHub
Evaluation & ObservabilityNo license specified

A Jev title-and-abstract screening experiment compared with Cohen Abstract Triage labels for an ADHD review.

WHERE JEV FITS

Asks Choice and Noul questions about eligibility, then combines them into include or exclude decisions.

Records metrics for specific dataset slices and question designs to examine screening errors.

11

GitHub
Evaluation & ObservabilityMIT

A Hermes goal-checking skill that saves requirements, creates a run prompt and checks completion evidence.

WHERE JEV FITS

Local checks handle machine-verifiable requirements; Jev assesses the remaining semantic conditions.

Connects goals with checkable criteria without replacing actual acceptance evidence with model opinions.

11

GitHub
Showing 24 / 29

About Awesome Jev

A community-maintained directory and radar for Jev, highlighting open-source projects with verified code and clear decision architectures.

What is Jev?

Jev is TypeSafe's decision model. This directory groups projects by how they use choices, scores and probability judgements, with implementation links to help you assess fit for your own task.

TypeSafe ↗

What gets listed?

We prioritize open-source projects with clear Jev implementation evidence. Repositories without verifiable source code are excluded from the main directory.

Directory data and design from Awesome Jev by logicrw, used under the MIT License.