Docs

Quickstart

Install JudgmentKit for your MCP client, then connect to the hosted Streamable HTTP endpoint.

curl -fsSL https://judgmentkit.ai/install | bash
curl -fsSL https://judgmentkit.ai/install | bash -s -- --client claude
curl -fsSL https://judgmentkit.ai/install | bash -s -- --client cursor

Codex is the default client. Use --client codex, --client claude, or --client cursor when scripting.

First 10 Minutes

Start with an interface task in your agent conversation. Describe who needs to do what, or point the agent at an existing interface and the task to improve.

Use JudgmentKit to build a signup form for a local workshop. Attendees pick an available session, enter their contact details, and receive a clear confirmation. Handle incomplete details and full sessions. Build and check the main task, then show what works and what remains unverified.

The agent should show a short working premise, build or revise the interface, try the important task, and make focused repairs. You can correct the premise as the work develops. The agent handles the review sequence.

Finish with the working interface, the important paths checked, and any missing verification. Contract acceptance is bounded to its checks; observed user task completion is separate evidence.

Replay the review mechanism

Use the replayable first-use fixture to see the AI-native design system as a contract loop, not a renderer. The fixture gives the agent one brief, one implementation contract input, one failing candidate, one repaired candidate, and the expected two-attempt transcript.

examples/ai-native-design-system/first-use.json
examples/ai-native-design-system/canonical-examples.json

Loop: create the implementation contract, review the failing candidate, read next_agent_action and grouped repair_instructions, repair the candidate, then resubmit and expect accept.

Canonical cases: setup/onboarding, operational dashboard, and high-stakes review/refund workflow. Each case includes the activity model, implementation contract input, failing candidate, repaired candidate, and proof expectation.

Runtime boundary: implementation_contract.design_system_source exposes the optional 17-contract React adapter candidate and its canonical registry. The root library, CLI, MCP, and visual_token_adapter remain framework-neutral. A complete design_system_adapter selects external_design_system; missing authorities fail instead of falling back to JudgmentKit defaults.

Planning Mode Examples

Use these examples to review whether an agent is using JudgmentKit well. A good planning response should make the activity, decision, outcome, and disclosure boundary clearer before it proposes UI structure.

Ready brief

Plan a UI for a support lead reviewing refund requests during daily triage. They decide whether each case is approved, sent to policy review, or returned for missing evidence. The outcome is a clear handoff with the next action and reason.

Good response: proceed to concept planning because the activity, participant, decision, and outcome are clear. Keep the plan centered on evidence review, decision options, and handoff.

Accept: approval, policy review, return for evidence, and handoff reasons are easy to compare and complete.

Reject: charts, widgets, or visual polish appear before the refund review work is named.

Vague brief

Plan a dashboard for the system.

Good response: infer and show the best provisional activity premise the prompt can support without inventing dashboard content. Ask at most one consequential question only when its answer would materially change the interaction and be costly to reverse.

Accept: the agent states its provisional premise and first direction, then asks the single highest-value question only if the unsupported activity, decision, or completion fork would change that direction.

Reject: a full dashboard plan with metrics, cards, charts, and navigation invented from no source context.

Implementation-heavy brief

Plan an admin UI from our JSON schema, database tables, tool call traces, prompt template, and API endpoints.

Good response: treat schemas, tables, traces, prompts, and endpoints as diagnostic details unless the task is explicitly setup, debugging, auditing, or integration work. Translate toward the user's activity before proposing a primary surface.

Accept: implementation terms move into diagnostics and the agent asks for the domain activity or decision behind the admin surface.

Reject: tables, schemas, prompt templates, tool calls, or API endpoints become the main product UI.

MCP

JudgmentKit supports MCP through the hosted Streamable HTTP endpoint at https://judgmentkit.ai/mcp. The installer registers that endpoint as judgmentkit in Codex, Claude Code, or Cursor. A browser GET to /mcp returns endpoint metadata; MCP clients should connect to the same URL with Streamable HTTP.

MCP tool responses include structuredContent as the stable machine-readable contract. Agents should translate it into ordinary domain language: a working premise, consequential decisions, the first direction, and at most one material question. Raw content[0].text is for explicit setup, audit, debugging, or integration work, not ordinary designer-facing conversation.

System Map

Use JudgmentKit before generation and across iterations. The agent coordinates the work; JudgmentKit returns reviewed contracts, evidence findings, and repair instructions.

This map describes the current source. Installed and hosted clients should check tools/list and tool input schemas for the capabilities available in their release.

JudgmentKit system design map An integration route showing caller-owned execution, library and MCP access, optional model proposals, deterministic activity and workflow review, implementation authority before handoff, frontend guidance, caller implementation and evidence, bounded implementation review, repair, and a separate observed human task result. Artifact Inspector remains review_required without trusted interactive attestation.
Outside the deterministic coreBuilder and agent
JudgmentKit kernelContracts and review
Outside the deterministic coreBuild, measure, and try
ready_for_review Packet repair / runtime retry; no attempt Blocked readiness Inspector: review_required repair_and_resubmit / stop_for_human UI and evidence repair Task findings
Client owns executionBuilder supplies the task.Agent calls tools, builds, and repairs.
Source brief + product contextCurrent brief and attributed facts.Resupply raw source at validatingboundaries; receipts are continuity.
Library and MCP accessMCP and library expose the review route.CLI: activity review + evidence preflight.Full packets or compact continuation.MCP wires bounded browser checks.
Optional model proposalsClient or explicit library helperinvokes an injected model caller.Activity/workflow candidates returnto deterministic review.
Agent resolves and resubmitsRepair packets before UI judgment.Repair failed UI after review.Changed premise: refresh sourceand dependent review packets.Human resolves escalated decisions.
1. Ground the activitycreate_activity_model_reviewreview_activity_model_candidateBaseline includes brief analysis.Negation and clause scope stay intact.Review external candidates before use.
2. Select the surfacerecommend_surface_typesrecommend_ui_workflow_profilesPurpose, confidence, origin, and conflicts.No evidence: review_required, no surface.Selection provenance is not authority.
3. Review the workflowreview_ui_workflow_candidatereview_cognitive_dimensions_candidateSource, actions, completion, disclosure.Cognitive Dimensions is optional;if supplied, it must be ready.
4. Set implementation authoritycreate_ui_implementation_contractDefault or complete external adapter.Required states and verification checks.Chart promise binds expected data.Incomplete external authority fails.
5. Gate the handoffcreate_ui_generation_handoffReady workflow + implementation contract.Revalidate current source and authority.Blocked readiness prevents generation.
6. Prepare frontend guidancecreate_frontend_generation_contextcreate_frontend_implementation_skill_contextReady handoff + resolved surface.Project context and verification plan.Portable instructions are optional.
7. Agent implements the interfaceUses the reviewed handoff and active authority.Client chooses the renderer; root is framework-neutral.React adapter is an optional candidate.
Run checks and collect evidenceAgent runs application/browser task checks.Trusted runtime measures eligible self-contained HTML.Chart labels, clipping, and selected data are observed.Expected data comes from the active contract.Static snapshots do not attest live transitions.
8. Preflight the evidencepreflight_ui_implementation_candidateShape and selectors before UI judgment.Packet repairs use no attempt.Runtime unavailable: retry_preflight.
9. Review implementationreview_ui_implementation_candidateContract, authority, states, evidence.Failed candidates are repair diagnostics.No automatic agent repair or publication.
Review verdictaccept / repair_and_resubmitstop_for_human / review_requiredAcceptance is bounded to checked claims.
Deferred Inspector verificationNo trusted interactive-attestation producer or verifier exists.Otherwise valid: review_required; next_agent_action: none.Primary artifact: external_not_reviewed. No automatic repair can close this limit.
Person tries the real taskObserve completion, confusion, and recovery.Record usefulness separately from contract acceptance.Task findings guide the next iteration.

Caller-owned execution: the builder’s agent uses the library or MCP for this workflow. The CLI supports activity analysis and review, plus preflight-implementation for evidence admission. MCP provides access and transport and wires supported browser checks. The agent owns inference, questions, generation, task checks, and repairs. Pass the exact current brief and attributed context through each validating boundary; integrity receipts establish continuity, not action authority.

Optional model proposals: an injected provider or the host agent may propose activity or workflow candidates. The caller sends the proposals to JudgmentKit for review; the deterministic kernel does not call a model.

Activity, surface, and workflow: activity review establishes the working premise, decisions, vocabulary, and disclosure rules. recommend_surface_types recommends among nine purposes: marketing, workbench, operator review, artifact inspector, form flow, dashboard monitor, content/report, setup/debug tool, and conversation. No positive surface evidence returns review_required with no recommended surface. Pattern confidence reflects core purpose evidence and competing purposes. Explicit selections identify caller, user, or agent origin; provenance does not grant action authority. Conflicts fail visibly. Workflow review checks grounding, supported actions, and completion or handoff.

Implementation contract before handoff: create_ui_implementation_contract defines approved primitives, required states, static checks, browser QA, and accessibility evidence. implementation_contract.design_system_source selects JudgmentKit defaults or a complete external adapter for tokens, fonts, icons, and components. Incomplete external adapters fail without falling back to JudgmentKit.

Ready handoff and frontend guidance: create_ui_generation_handoff requires a ready workflow review and a valid implementation contract. An optional Cognitive Dimensions review blocks handoff when supplied and not ready. create_frontend_generation_context combines the ready handoff, selected surface type, and frontend context. create_frontend_implementation_skill_context compiles portable implementation guidance. The client builds and runs the interface; renderer choice follows the active design-system source.

Evidence admission before review: call preflight_ui_implementation_candidate to check evidence structure and declared selectors. ready_for_review admits the packet to substantive review. repair_evidence_packet or retry_evidence_preflight performs no substantive review and consumes no implementation attempt. A valid packet describing a failing interface still fails implementation review.

Implementation evidence and review: the caller supplies static, state, accessibility, and browser QA evidence to review_ui_implementation_candidate. Eligible self-contained HTML can receive trusted visual-composition and chart observations through the MCP route. Chart checks measure label collisions, clipping, and selected-data correspondence at the required states and viewports, using expected data attributed to the active contract. Those supported observations do not authenticate data truth or attest live transitions. The result directs the agent to accept, repair_and_resubmit, or stop_for_human, or keeps the implementation review_required when an authority requirement remains unresolved.

Full and compact packets: full output is the default. Optional packet_format: "compact" carries readable active guidance and a bounded lossless continuation. Pass the complete envelope to downstream MCP tools; library clients expand it with the packet helpers first. Raw source is still required. For large implementation evidence, the local judgmentkit/packets helpers prepare candidate continuations without dropping snapshots or declared evidence.

Artifact Inspector limit: this proposed profile separates JudgmentKit-owned chrome and overlays from the external artifact. JudgmentKit has no trusted interactive-attestation producer or verifier, so an otherwise valid Inspector implementation remains review_required. Static browser measurements cannot close that requirement.

Caller-owned iteration: the agent repairs the implementation and resubmits evidence, or stops for human help when the attempt policy requires it. Changes to source decisions require fresh affected reviews using the updated brief and attributed context.

Human task outcome: an intended user’s observed task completion and understanding remain separate from implementation acceptance. A passing contract cannot establish usefulness by itself.

Separate presentation tools: create_slide_deck plans JudgmentKit presentation-theme decks from slide content. Hosted callers can use dry-run planning; PPTX export requires a local artifact runtime. These tools are separate from the UI generation path.

If a review cannot proceed, the agent explains what is missing and resolves it using available evidence. It asks you when a product decision or governing policy is needed.

Activity Review

Call create_activity_model_review before generating UI from a brief. Treat its deterministic candidate as a baseline, let the host model infer the complete best-current activity case, then call review_activity_model_candidate before trusting that inferred case.

Workflow Review

Call review_ui_workflow_candidate before accepting an agent-proposed workflow. It checks source grounding, action support, completion or handoff clarity, and leakage containment.

Cognitive Dimensions Review

Call review_cognitive_dimensions_candidate when a workflow or implementation candidate needs review for domain mapping, evidence near action, hidden dependencies, premature commitment, progressive evaluation, change cost, memory-heavy transitions, or disclosure leakage. Findings are diagnostic guidance for agents and reviewers; do not copy Cognitive Dimensions terminology into product UI.

Surface Type

Call recommend_surface_types after activity review and before workflow or frontend implementation guidance. Surface type is activity-purpose guidance, not a visual theme.

Marketing surface
Persuade, orient, convert, or explain an offer.
Workbench
Help a user repeatedly inspect, compare, decide, and act.
Operator review
Review AI- or system-produced work, evidence, risk, and handoff.
Artifact inspector
Inspect one rendered artifact in place, act on a semantic locus, and leave an explicit artifact-local result.
Form flow
Collect or change structured information with validation.
Dashboard monitor
Track status, exceptions, trends, or operational health.
Content or report
Read, understand, cite, or share information.
Setup or debugging tool
Configure, inspect, test, or troubleshoot machinery.
Conversation
Support open-ended exchange where the thread is the product surface.

Status: proposed

Artifact Inspector

Artifact Inspector is a proposed interaction model for work centered on one rendered artifact. Use it only when the artifact must remain visible and primary, the person must select a semantic locus within it, and supporting evidence, actions, or results are meaningful in relation to that locus.

Keep the existing surface type when a queue, case, collection, report, dashboard, form, setup flow, or conversation is primary. A live artifact keeps its native behavior until the person explicitly enters inspection mode.

Identifiers and current status

surface_type: artifact_inspector
workflow_profile: artifact-inspector-ui
frontend_surface_profile: judgmentkit.artifact-inspector.v1
topology_kind: artifact_centered
status: proposed
implementation_review_status: review_required
primary_artifact_review_status: external_not_reviewed

Authority boundary

JudgmentKit governs the inspector chrome and inspection overlay, not the artifact itself. The declared external authority owns the artifact’s typography, components, color, elevation, internal layout, semantics, and native interactions. Reviews report chrome, overlay, artifact preservation, and boundary behavior separately.

Current review boundary

JudgmentKit 0.8.0 can identify and validate the artifact-centered contract and carry it into generation guidance. It does not yet have a trusted interactive-attestation producer or verifier, so an otherwise valid implementation remains review_required.

Screenshots, static metadata, caller-authored evidence, or unchanged fingerprints cannot close that gate. Future acceptance must verify real pointer, touch, keyboard, and assistive-technology crossings; focus order and return; overlay obstruction and target drift; style isolation in both directions; artifact preservation; and required states across desktop and narrow viewports.

Handoff

Call create_ui_generation_handoff only on a ready workflow review, resupplying the exact current brief and attributed context_items so protected risk and workflow authority are revalidated from raw source. If the gate blocks, resolve the material ambiguity or authoritative-source boundary first.

Implementation Contract

Call create_ui_implementation_contract before final handoff so generated UI has approved primitives, state coverage, implementation_contract.design_system_source, implementation_contract.local_component_authority, implementation_contract.visual_token_adapter, implementation_contract.default_ai_native_design_system, static checks, browser QA expectations, implementation_contract.visual_asset_policy, and implementation_contract.accessibility_policy. Call review_ui_implementation_candidate before accepting generated UI code or evidence. Visual-heavy pages need browser-rendered contrast/readability evidence for text over images, canvas, WebGL, video, gradients, or generated visuals.

Frontend Context

Call create_frontend_generation_context after the handoff gate when an agent needs frontend implementation guidance with selected surface type, project context, and verification expectations. Resupply the exact current brief and attributed context_items. Call create_frontend_implementation_skill_context with that same raw source and the ready frontend context when an MCP client needs compiled implementation guidance instead of repo-local skill access.

Slide Decks

Call create_slide_deck when an allowed brief, workflow review, handoff, or implementation evidence should become a JudgmentKit presentation, PowerPoint, or PPTX. The tool returns selected templates and content keys in dry-run mode, and writes PPTX artifacts only from a local @oai/artifact-tool runtime under the guarded output directory.

Guidance Profiles

Call recommend_ui_workflow_profiles when a brief sounds like specialized review work. Pass profile_id: "operator-review-ui" only when the recommendation evidence supports it. Artifact-centered work may use the proposed artifact-inspector-ui workflow profile with judgmentkit.artifact-inspector.v1; that profile guides generation but does not change its review_required status.