Docs
Quickstart
Install JudgmentKit for your MCP client, then connect to the hosted Streamable HTTP endpoint.
curl -fsSL https://judgmentkit.ai/install | bash
curl -fsSL https://judgmentkit.ai/install | bash -s -- --client claude
curl -fsSL https://judgmentkit.ai/install | bash -s -- --client cursor
Codex is the default client. Use --client codex, --client claude, or --client cursor when scripting.
First 10 Minutes
Start with an interface task in your agent conversation. Describe who needs to do what, or point the agent at an existing interface and the task to improve.
Use JudgmentKit to build a signup form for a local workshop. Attendees pick an available session, enter their contact details, and receive a clear confirmation. Handle incomplete details and full sessions. Build and check the main task, then show what works and what remains unverified.
The agent should show a short working premise, build or revise the interface, try the important task, and make focused repairs. You can correct the premise as the work develops. The agent handles the review sequence.
Finish with the working interface, the important paths checked, and any missing verification. Contract acceptance is bounded to its checks; observed user task completion is separate evidence.
Replay the review mechanism
Use the replayable first-use fixture to see the AI-native design system as a contract loop, not a renderer. The fixture gives the agent one brief, one implementation contract input, one failing candidate, one repaired candidate, and the expected two-attempt transcript.
examples/ai-native-design-system/first-use.json
examples/ai-native-design-system/canonical-examples.json
Loop: create the implementation contract, review the failing candidate, read next_agent_action and grouped repair_instructions, repair the candidate, then resubmit and expect accept.
Canonical cases: setup/onboarding, operational dashboard, and high-stakes review/refund workflow. Each case includes the activity model, implementation contract input, failing candidate, repaired candidate, and proof expectation.
Runtime boundary: implementation_contract.design_system_source exposes the optional 17-contract React adapter candidate and its canonical registry. The root library, CLI, MCP, and visual_token_adapter remain framework-neutral. A complete design_system_adapter selects external_design_system; missing authorities fail instead of falling back to JudgmentKit defaults.
Planning Mode Examples
Use these examples to review whether an agent is using JudgmentKit well. A good planning response should make the activity, decision, outcome, and disclosure boundary clearer before it proposes UI structure.
Ready brief
Plan a UI for a support lead reviewing refund requests during daily triage. They decide whether each case is approved, sent to policy review, or returned for missing evidence. The outcome is a clear handoff with the next action and reason.
Good response: proceed to concept planning because the activity, participant, decision, and outcome are clear. Keep the plan centered on evidence review, decision options, and handoff.
Accept: approval, policy review, return for evidence, and handoff reasons are easy to compare and complete.
Reject: charts, widgets, or visual polish appear before the refund review work is named.
Vague brief
Plan a dashboard for the system.
Good response: infer and show the best provisional activity premise the prompt can support without inventing dashboard content. Ask at most one consequential question only when its answer would materially change the interaction and be costly to reverse.
Accept: the agent states its provisional premise and first direction, then asks the single highest-value question only if the unsupported activity, decision, or completion fork would change that direction.
Reject: a full dashboard plan with metrics, cards, charts, and navigation invented from no source context.
Implementation-heavy brief
Plan an admin UI from our JSON schema, database tables, tool call traces, prompt template, and API endpoints.
Good response: treat schemas, tables, traces, prompts, and endpoints as diagnostic details unless the task is explicitly setup, debugging, auditing, or integration work. Translate toward the user's activity before proposing a primary surface.
Accept: implementation terms move into diagnostics and the agent asks for the domain activity or decision behind the admin surface.
Reject: tables, schemas, prompt templates, tool calls, or API endpoints become the main product UI.
MCP
JudgmentKit supports MCP through the hosted Streamable HTTP endpoint at https://judgmentkit.ai/mcp. The installer registers that endpoint as judgmentkit in Codex, Claude Code, or Cursor. A browser GET to /mcp returns endpoint metadata; MCP clients should connect to the same URL with Streamable HTTP.
MCP tool responses include structuredContent as the stable machine-readable contract. Agents should translate it into ordinary domain language: a working premise, consequential decisions, the first direction, and at most one material question. Raw content[0].text is for explicit setup, audit, debugging, or integration work, not ordinary designer-facing conversation.
System Map
Use JudgmentKit before generation and across iterations. The agent coordinates the work; JudgmentKit returns reviewed contracts, evidence findings, and repair instructions.
This map describes the current source. Installed and hosted clients should check tools/list and tool input schemas for the capabilities available in their release.
Caller-owned execution: the builder’s agent uses the library or MCP for this workflow. The CLI supports activity analysis and review, plus preflight-implementation for evidence admission. MCP provides access and transport and wires supported browser checks. The agent owns inference, questions, generation, task checks, and repairs. Pass the exact current brief and attributed context through each validating boundary; integrity receipts establish continuity, not action authority.
Optional model proposals: an injected provider or the host agent may propose activity or workflow candidates. The caller sends the proposals to JudgmentKit for review; the deterministic kernel does not call a model.
Activity, surface, and workflow: activity review establishes the working premise, decisions, vocabulary, and disclosure rules. recommend_surface_types recommends among nine purposes: marketing, workbench, operator review, artifact inspector, form flow, dashboard monitor, content/report, setup/debug tool, and conversation. No positive surface evidence returns review_required with no recommended surface. Pattern confidence reflects core purpose evidence and competing purposes. Explicit selections identify caller, user, or agent origin; provenance does not grant action authority. Conflicts fail visibly. Workflow review checks grounding, supported actions, and completion or handoff.
Implementation contract before handoff: create_ui_implementation_contract defines approved primitives, required states, static checks, browser QA, and accessibility evidence. implementation_contract.design_system_source selects JudgmentKit defaults or a complete external adapter for tokens, fonts, icons, and components. Incomplete external adapters fail without falling back to JudgmentKit.
Ready handoff and frontend guidance: create_ui_generation_handoff requires a ready workflow review and a valid implementation contract. An optional Cognitive Dimensions review blocks handoff when supplied and not ready. create_frontend_generation_context combines the ready handoff, selected surface type, and frontend context. create_frontend_implementation_skill_context compiles portable implementation guidance. The client builds and runs the interface; renderer choice follows the active design-system source.
Evidence admission before review: call preflight_ui_implementation_candidate to check evidence structure and declared selectors. ready_for_review admits the packet to substantive review. repair_evidence_packet or retry_evidence_preflight performs no substantive review and consumes no implementation attempt. A valid packet describing a failing interface still fails implementation review.
Implementation evidence and review: the caller supplies static, state, accessibility, and browser QA evidence to review_ui_implementation_candidate. Eligible self-contained HTML can receive trusted visual-composition and chart observations through the MCP route. Chart checks measure label collisions, clipping, and selected-data correspondence at the required states and viewports, using expected data attributed to the active contract. Those supported observations do not authenticate data truth or attest live transitions. The result directs the agent to accept, repair_and_resubmit, or stop_for_human, or keeps the implementation review_required when an authority requirement remains unresolved.
Full and compact packets: full output is the default. Optional packet_format: "compact" carries readable active guidance and a bounded lossless continuation. Pass the complete envelope to downstream MCP tools; library clients expand it with the packet helpers first. Raw source is still required. For large implementation evidence, the local judgmentkit/packets helpers prepare candidate continuations without dropping snapshots or declared evidence.
Artifact Inspector limit: this proposed profile separates JudgmentKit-owned chrome and overlays from the external artifact. JudgmentKit has no trusted interactive-attestation producer or verifier, so an otherwise valid Inspector implementation remains review_required. Static browser measurements cannot close that requirement.
Caller-owned iteration: the agent repairs the implementation and resubmits evidence, or stops for human help when the attempt policy requires it. Changes to source decisions require fresh affected reviews using the updated brief and attributed context.
Human task outcome: an intended user’s observed task completion and understanding remain separate from implementation acceptance. A passing contract cannot establish usefulness by itself.
Separate presentation tools: create_slide_deck plans JudgmentKit presentation-theme decks from slide content. Hosted callers can use dry-run planning; PPTX export requires a local artifact runtime. These tools are separate from the UI generation path.
If a review cannot proceed, the agent explains what is missing and resolves it using available evidence. It asks you when a product decision or governing policy is needed.
Activity Review
Call create_activity_model_review before generating UI from a brief. Treat its deterministic candidate as a baseline, let the host model infer the complete best-current activity case, then call review_activity_model_candidate before trusting that inferred case.
Workflow Review
Call review_ui_workflow_candidate before accepting an agent-proposed workflow. It checks source grounding, action support, completion or handoff clarity, and leakage containment.
Cognitive Dimensions Review
Call review_cognitive_dimensions_candidate when a workflow or implementation candidate needs review for domain mapping, evidence near action, hidden dependencies, premature commitment, progressive evaluation, change cost, memory-heavy transitions, or disclosure leakage. Findings are diagnostic guidance for agents and reviewers; do not copy Cognitive Dimensions terminology into product UI.
Surface Type
Call recommend_surface_types after activity review and before workflow or frontend implementation guidance. Surface type is activity-purpose guidance, not a visual theme.
- Marketing surface
- Persuade, orient, convert, or explain an offer.
- Workbench
- Help a user repeatedly inspect, compare, decide, and act.
- Operator review
- Review AI- or system-produced work, evidence, risk, and handoff.
- Artifact inspector
- Inspect one rendered artifact in place, act on a semantic locus, and leave an explicit artifact-local result.
- Form flow
- Collect or change structured information with validation.
- Dashboard monitor
- Track status, exceptions, trends, or operational health.
- Content or report
- Read, understand, cite, or share information.
- Setup or debugging tool
- Configure, inspect, test, or troubleshoot machinery.
- Conversation
- Support open-ended exchange where the thread is the product surface.
Status: proposed
Artifact Inspector
Artifact Inspector is a proposed interaction model for work centered on one rendered artifact. Use it only when the artifact must remain visible and primary, the person must select a semantic locus within it, and supporting evidence, actions, or results are meaningful in relation to that locus.
Keep the existing surface type when a queue, case, collection, report, dashboard, form, setup flow, or conversation is primary. A live artifact keeps its native behavior until the person explicitly enters inspection mode.
Identifiers and current status
surface_type: artifact_inspector
workflow_profile: artifact-inspector-ui
frontend_surface_profile: judgmentkit.artifact-inspector.v1
topology_kind: artifact_centered
status: proposed
implementation_review_status: review_required
primary_artifact_review_status: external_not_reviewed
Authority boundary
JudgmentKit governs the inspector chrome and inspection overlay, not the artifact itself. The declared external authority owns the artifact’s typography, components, color, elevation, internal layout, semantics, and native interactions. Reviews report chrome, overlay, artifact preservation, and boundary behavior separately.
Current review boundary
JudgmentKit 0.8.0 can identify and validate the artifact-centered contract and carry it into generation guidance. It does not yet have a trusted interactive-attestation producer or verifier, so an otherwise valid implementation remains review_required.
Screenshots, static metadata, caller-authored evidence, or unchanged fingerprints cannot close that gate. Future acceptance must verify real pointer, touch, keyboard, and assistive-technology crossings; focus order and return; overlay obstruction and target drift; style isolation in both directions; artifact preservation; and required states across desktop and narrow viewports.
Handoff
Call create_ui_generation_handoff only on a ready workflow review, resupplying the exact current brief and attributed context_items so protected risk and workflow authority are revalidated from raw source. If the gate blocks, resolve the material ambiguity or authoritative-source boundary first.
Implementation Contract
Call create_ui_implementation_contract before final handoff so generated UI has approved primitives, state coverage, implementation_contract.design_system_source, implementation_contract.local_component_authority, implementation_contract.visual_token_adapter, implementation_contract.default_ai_native_design_system, static checks, browser QA expectations, implementation_contract.visual_asset_policy, and implementation_contract.accessibility_policy. Call review_ui_implementation_candidate before accepting generated UI code or evidence. Visual-heavy pages need browser-rendered contrast/readability evidence for text over images, canvas, WebGL, video, gradients, or generated visuals.
Frontend Context
Call create_frontend_generation_context after the handoff gate when an agent needs frontend implementation guidance with selected surface type, project context, and verification expectations. Resupply the exact current brief and attributed context_items. Call create_frontend_implementation_skill_context with that same raw source and the ready frontend context when an MCP client needs compiled implementation guidance instead of repo-local skill access.
Slide Decks
Call create_slide_deck when an allowed brief, workflow review, handoff, or implementation evidence should become a JudgmentKit presentation, PowerPoint, or PPTX. The tool returns selected templates and content keys in dry-run mode, and writes PPTX artifacts only from a local @oai/artifact-tool runtime under the guarded output directory.
Guidance Profiles
Call recommend_ui_workflow_profiles when a brief sounds like specialized review work. Pass profile_id: "operator-review-ui" only when the recommendation evidence supports it. Artifact-centered work may use the proposed artifact-inspector-ui workflow profile with judgmentkit.artifact-inspector.v1; that profile guides generation but does not change its review_required status.