Search: judge
48 result(s) on page 1
Judge
A comprehensive AI agent skill for navigating court systems and legal proceedings. Helps you understand what type of court handles your situation, prepares y...
quarantinedcourthearingjudgeClawHubjudge
Indexed by skills.sh from neolabhq/context-engineering-kit
quarantinedskills.sh- Registry
- skills.sh
- Category
- Domain· inferred
- Version
- 1.0.0
- Author
- neolabhq
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__judgeLLM_as_Judge
Cross-model verification for complex tasks. Spawn a judge subagent with a different model to review plans, code, architecture, or decisions before execution....
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__LLM_as_JudgeStrict_Paper_Judge
Strictly judge whether a research paper is worth following, reading, or recommending. Use for paper triage, paper reviews, literature evaluation, arXiv scree...
quarantinedClawHub- Registry
- ClawHub
- Category
- Research· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Strict_Paper_JudgeAmazon_Listing_Judge
Grade Amazon product listing quality. Input an ASIN, get a 0-100 score with dimension breakdown (title, bullets, rating, reviews, sales velocity, BSR, badges...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Amazon_Listing_JudgeJudge_Human
Vote and submit AI evaluation signals on ethical, cultural, and content stories alongside human crowds. Includes an autonomous heartbeat orchestrator (heartb...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Judge_HumanLlm_As_Judge
Build a cost-efficient LLM evaluation ensemble with sampling, tiebreakers, and deterministic validators. Learned from 600+ production runs judging local Olla...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Llm_As_JudgeLlm_Judge
Use when comparing two or more code implementations against a spec or requirements doc. Triggers on "which repo is better", "compare these implementations",...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Llm_Judgedo-and-judge
Indexed by skills.sh from neolabhq/context-engineering-kit
quarantinedskills.sh- Registry
- skills.sh
- Category
- Domain· inferred
- Version
- 1.0.0
- Author
- neolabhq
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__do-and-judgejudge-with-debate
Indexed by skills.sh from neolabhq/context-engineering-kit
quarantinedskills.sh- Registry
- skills.sh
- Version
- 1.0.0
- Author
- neolabhq
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__judge-with-debateskill-judge
Indexed by skills.sh from davila7/claude-code-templates
quarantinedskills.sh- Registry
- skills.sh
- Category
- Domain· inferred
- Version
- 1.0.5
- Author
- davila7
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__skill-judgewrite-judge-prompt
Indexed by skills.sh from hamelsmu/evals-skills
quarantinedskills.sh- Registry
- skills.sh
- Version
- 1.0.0
- Author
- hamelsmu
curl -s /v1/skills/community__axehub:skills.sh__skills.sh__write-judge-promptGenerate_Judgements
Use when creating or updating test judgement definitions (judge_definitions) for an agent skill evaluation YAML config. Analyzes a skill's SKILL.md and refer...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Generate_JudgementsJEP_Primitive_Skills
JEP Primitive Skills — Atomic Reference Implementations of Judge, Delegate, Terminate, Verify for Agent Collaboration Grammar
quarantinedatomicdelegategrammarClawHubCodex_Handoff_(OpenClaw_Plans,_Codex_Codex,_OpenClaw_Judges)
Offload finalized coding plans to Codex CLI for automated execution. Use when: user says "hand off to codex", "let codex do it", "offload to codex", runs /co...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Codex_Handoff_(OpenClaw_Plans,_Codex_Codex,_OpenClaw_Judges)Fetch_personal_academic_records_from_global_databases_via_individual_ID_lookup_queries._Extract_complete_educational_profiles_covering_attended_institutions,awarded_degrees,_majors_and_GPA_metrics.Recruiters,_HR_teams_and_hiring_managers_validate_applicant_education_histories,_assess_candidate_credentials_anddeliver_data-backed_hiring_judgements._Streamline_pre-employment_screening,_formal_background_verification_and_end-to-end_talent_assessment_workflows_forcorporate_recruitment_teams.
Verify overseas candidates’ education history including degrees, majors and GPA. Help HR teams run background checks and select qualified applicants for hiri...
quarantinedHRacademicbackgroundClawHub- Registry
- ClawHub
- Category
- Research
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Fetch_personal_academic_records_from_global_databases_via_individual_ID_lookup_queries._Extract_complete_educational_profiles_covering_attended_institutions,awarded_degrees,_majors_and_GPA_metrics.Recruiters,_HR_teams_and_hiring_managers_validate_applicant_education_histories,_assess_candidate_credentials_anddeliver_data-backed_hiring_judgements._Streamline_pre-employment_screening,_formal_background_verification_and_end-to-end_talent_assessment_workflows_forcorporate_recruitment_teams.Refinement_Loop_(Opus_4.8_Edition)
Design and run iterative generate→critique→revise loops optimized for Claude Opus 4.8, with thinking-as-critic, cost controls, and model routing.
quarantinedcritiqueiterationllm-as-judgeClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Refinement_Loop_(Opus_4.8_Edition)blackbox
Delegate coding tasks to the Blackbox AI multi-model CLI.
quarantinedlinuxmacoswindowsoptional- Registry
- optional
- Category
- AI Agents
- Version
- 1.0.0
- Author
- Hermes Agent (Nous Research)
- License
- MIT
curl -s /v1/skills/community__axehub:optional__optional__blackboxeval-observability
verifiedFirst-party- Registry
- First-party
- Category
- Evaluation
- Version
- 1.0.0
- Author
- AXE
- License
- MIT
curl -s /v1/skills/eval-observabilityAdopt_A_Gustowl
Nocturnal and judgmental. Anthropic called it a Gustowl. We called it an Owl. Both judge you silently. Real-time hunger. Permanent death. 5 evolution stages....
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Adopt_A_GustowlAdvanced_Evaluation
This skill should be used when the user asks to "implement LLM-as-judge", "compare model outputs", "create evaluation rubrics", "mitigate evaluation bias", o...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Advanced_EvaluationAgent_Scorecard
Configurable quality evaluation for AI agent outputs. Define criteria, run evaluations, track quality over time. No LLM-as-judge, no API calls, pattern-based...
quarantinedagentevaluationmetricsClawHub- Registry
- ClawHub
- Category
- AI Agents
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Agent_ScorecardBlackbox
Delegate coding tasks to Blackbox AI CLI agent. Multi-model agent with built-in judge that runs tasks through multiple LLMs and picks the best result. Requir...
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__BlackboxCase_Writer_Hybrid
Expand a structured brief in `content-production/inbox/` into a reusable long-form markdown article draft, then run a local writer / critic / judge quality l...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Case_Writer_HybridCritical_Debater_Suite
Multi-agent adversarial debate system with 4 roles (Pro, Con, Judge, Orchestrator), per-round evidence refresh, 5-element reasoning chains, and structured bi...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Critical_Debater_SuiteCrypto_One_Way_Market
Fetch cryptocurrency OHLCV candle data and judge whether the market is in a one-way bullish or bearish trend. Use when the user asks to pull crypto market da...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Crypto_One_Way_MarketDark_Factory_Skill
Manage multiple SaaS startups simultaneously with CEO-driven orchestration, product agents, ChatDev code generation, and a 3-Gate BUILD, TEST, JUDGE pipeline.
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Dark_Factory_SkillDebate_Learning_Workflow
Run evidence-backed multi-agent debates (A/B/Opponent3/Judge) with 20-40 rounds, loophole analysis, and universal actionable lesson extraction.
quarantineddebateworkflowClawHub- Registry
- ClawHub
- Category
- Productivity
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Debate_Learning_WorkflowDecision_Autopsy
Judge a past decision by its PROCESS, not its outcome — because good decisions lose and bad decisions win, and teams that can't tell the difference learn the...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Decision_AutopsyDidi
Help users make better Didi ride-booking decisions from public ride-choice logic. Use when the user wants to compare ride types, judge whether a ride option...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__DidiEmail_Importance_Content_Analysis
Judge whether an email is important/urgent using content-based analysis rather than sender name or mailbox labels (which can be spoofed). Use when asked to triage emails, decide priority, detect phishing/social-engineering, or recommend next actions (reply/pay/login/download/c…
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Email_Importance_Content_AnalysisEval_Rubric_Designer
Design a scoring rubric and LLM-as-judge prompt to evaluate the quality of an AI feature's output. Use when asked to create an eval rubric, define quality di...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Eval_Rubric_DesignerFeishu_Ingest
Poll Feishu groups for new messages, download message resources, read Feishu docs/wiki/sheets/bitables, compile message-source markdown, judge which chat segments are worth preserving, and ingest valuable materials into the Research KB through a prepare/apply workflow.
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Feishu_IngestGhosted_Skill
Ghosted is an AI coach for people dealing with ghosting, being left on read, or watching a once-warm chat suddenly go cold. It helps you judge whether to fol...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Ghosted_SkillLlm_Eval_Router
Shadow-test local Ollama models against a cloud baseline with a multi-judge ensemble. Automatically promotes models when statistically proven equivalent — re...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Llm_Eval_RouterLora_Finetune
LoRA fine-tuning pipeline for Stable Diffusion on Apple Silicon — dataset prep, training, evaluation with LLM-as-judge scoring. Use when fine-tuning image ge...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Lora_FinetuneMoot_Court_AI
Simulate a full Chinese civil court hearing with 4 role-based agents (clerk, plaintiff, defendant, judge) orchestrated by deterministic Lobster workflow.
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Moot_Court_AIOpenclaw_Auto_Training_Skill
Autonomous QA evaluation loop — runs domain-specific tasks against yourself, scores responses with an LLM judge, installs missing skills, and logs knowledge...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Openclaw_Auto_Training_SkillPeer_Reviewer
AI-powered academic paper reviewer. Uses a multi-agent system (Deconstructor, Devil's Advocate, Judge) to analyze papers for logical flaws, contradictions, and empirical validity.
quarantinedClawHub- Registry
- ClawHub
- Category
- Research· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Peer_ReviewerThe_Arena_—_AI_Debate_Moderator
Turn a Discord server into a moderated debate arena with an AI judge. Supports multiple debate formats, configurable personas, scored verdicts, and a persist...
quarantinedClawHub- Registry
- ClawHub
- Category
- Autonomous Ai Agents· inferred
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__The_Arena_—_AI_Debate_ModeratorTinkerClaw_Memory_Bench
Be one of the first to benchmark your agent's memory — and help shape how AI remembers. Runs a peer-review-grade evaluation suite (LLM-as-judge, nDCG/MAP/MRR...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__TinkerClaw_Memory_BenchWei_Cross_Research
Cross-validate research answers by querying multiple LLMs in parallel with judge-based synthesis. Reduces hallucination and surfaces model disagreements for...
quarantinedconsensusmulti-modelreasoningClawHub- Registry
- ClawHub
- Category
- Research
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__Wei_Cross_Researchdayday
English skill based on the public product messaging of MeiRiYiLian. Use when the user wants to understand the product, judge whether it fits a learning goal,...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__daydayelm
Help users make better Eleme ordering decisions from public merchant and promotion information. Use when the user wants to compare Eleme stores, judge whethe...
quarantinedassisted-orderingelemeelmClawHub- Registry
- ClawHub
- Category
- Business & Finance
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__elm内容策略诊断
Content-strategy diagnosis for Chinese creator work. Use when Codex needs to judge whether a topic, draft, or content plan should become a short video, post,...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__内容策略诊断对标筛选
Benchmark filtering for Chinese creator, OPC, and one-person-business work. Use when Codex needs to judge whether a person, creator, or business is actually...
quarantinedClawHub- Registry
- ClawHub
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:ClawHub__ClawHub__对标筛选search-cases
Search the NY State Unified Court System eCourts dockets (WebCivil Supreme/Local, WebCriminal, WebFamily) by index/docket number, party, attorney, or judge and return matching cases as structured JSON. Read-only.
quarantinedlegalcourt-recordsdocketsbrowse.sh- Registry
- browse.sh
- Category
- Legal
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:browse.sh__browse.sh__search-casessearch-supreme-court-judgments
Search the BC Supreme Court (bccourts.ca) public judgments index by full-text (boolean), case name, neutral citation, judge, docket, registry, or date range. First-class support for landlord-tenant matters that reached BCSC via judicial review of Residential Tenancy Branch dec…
quarantinedlegalcase-lawcourtbrowse.sh- Registry
- browse.sh
- Category
- Legal
- Version
- 1.0.0
curl -s /v1/skills/community__axehub:browse.sh__browse.sh__search-supreme-court-judgments