Productivity
evaluation-framework
Provides weighted scoring, rubrics, and decision-threshold patterns
Business
analizar-licitacion-osce
Evalúa una propuesta técnica de un concurso OSCE (supervisión/consultoría de obra) leyendo bases.pdf + propuesta.pdf. Orquesta subagentes — agent-bases (requisitos/factores/persona…
Testing
pulser
Diagnose and test Claude Code skills against Anthropic's 7 principles. Scans SKILL.md files, checks 8 rules (gotchas, description, allowed-tools, file-size, structure, frontmatter,…
AI / ML
darwin-skill
Darwin Skill (达尔文.skill): autonomous skill optimizer inspired by Karpathy's autoresearch. Evaluates SKILL.md files using an 8-dimension rubric (structure + effectiveness), runs hil…
AI / ML
empirical-prompt-tuning
Fetch and execute mizchi's empirical-prompt-tuning skill at runtime. Use when evaluating or iteratively refining an agent-facing prompt (skill / slash command / task prompt / CLAUD…
Research
research-idea-benchmark
This skill should be used when the user asks to "evaluate my research idea", "benchmark my research", "score my paper idea", "assess research quality", "check if my idea is publish…
AI / ML
metaharness
Design, evaluate, and iterate on Claude Code agents, commands, and skills with structural best practices. Use when creating a new .claude/agents/ file, updating an existing agent's…
Research
assess-outline
Evaluates research paper outlines against empirical research criteria for research question clarity, contributions, hypotheses, data description, results approach, and robustness c…
Testing
bench-build
Builds benchmark tasks from pending backlog entries by delegating each to a subagent that fully authors the task (pin, overlay, prompt, assertions, weights, self-check) and produce…
AI / ML
forge-evals
Design evaluations for LLM features including golden datasets, rubric scoring, LLM-as-judge calibration, CI regression detection, online A/B tests, cost and latency budgets, and ad…
AI / ML
ce-optimize
Run metric-driven iterative optimization loops. Define a goal, add measurement scaffolding, execute parallel experiments across approaches, score results against gates or quality j…
Research
cjc-experiments
Guides experimental design, evaluation, and statistical reporting for long-form computer science submissions, ensuring evidence matches contribution type, baselines are fair, datas…
Content
通用-平台签约评估框架
Evaluates creative works for platform acceptance or serialization potential using strict, conservative criteria that prioritize negative evidence. Provides a unified assessment fra…
Testing
loop-evals
Designs the evaluation harness for agent loops, establishing trustworthy verification through a 7-layer suite with false-completion-rate and repair-productivity as core metrics. En…
Testing
agento11y-test-starter
Builds an offline test suite for an agent before deployment by analyzing its code, prompts, and tools to generate labeled cases (happy/edge/adversarial) with scoring recommendation…
AI / ML
mllm-eval
Designs or audits a model-agnostic evaluation harness for LLMs or multimodal models on clinical tasks, covering reference standards, clinical metrics, faithfulness checks, contamin…
AI / ML
experiment-lab
Run reproducible CS/AI experiments, model evaluation, regression/classification/clustering analyses, bioinformatics workflows, QC, differential expression, single-cell starter anal…
AI / ML
ai-engineering
Authoritative reference on AI tooling selection and architecture decisions. Consult when evaluating frameworks, memory systems, RAG pipelines, serving infrastructure, or evaluation…
Productivity
legw-rmap-evaluierung-und-aenderung
Manages the full lifecycle of a rulemap norm: versioning, drag-and-drop edits in the builder, NKRG/GGO evaluation, impact monitoring, and feedback from implementation. Outputs chan…
AI / ML
machine-learning-mitchell
Applies Mitchell's machine learning framework—well-posed problems, hypothesis space, inductive bias, overfitting, bias-variance tradeoff, proper evaluation with statistical signifi…
Content
deductive-method
Apply the Court of Master Sommeliers Deductive Tasting Grid to systematically evaluate any wine. Use when describing a wine from scratch, practicing blind tasting, or building a st…
Testing
os-eval-runner
Stateless evaluation engine that scores and gates skill improvement iterations using headless Python scripts. Use for "evaluate this skill", "run autoresearch loop", "optimize this…
AI / ML
agento11y
Inspects and manages Grafana Agent observability resources—conversations, generations, evaluators, rules, scores, and templates—via gcx. Use for listing or searching conversations,…
Testing
gitlearnos-review
Recognize and write back a learner answer or attempted solution, then administer, evaluate, and schedule an existing GitLearnOS question set using observable evidence. Use even whe…
AI / ML
ai-feature-eval-harness
Designs an evaluation plan for AI product features with measurable criteria, held-out datasets, code- and LLM-based grading tiers, pass thresholds, and AI_EVAL_PLAN.md output. Appl…
Data
data-source-eval
Evaluates new data sources for feasibility, integration effort, and expected value. Runs a standardized spike covering objectives, API discovery, technical viability, architecture,…
AI / ML
devlab-eval-driven-agent
An evaluation-driven methodology for building production AI agents, centered on curated datasets, mock isolation, standardized comparison, automated scoring scripts, and regression…
AI / ML
df-review-reasoning-layer
Reviews layered prompt systems that assemble dialectical reasoning across apps, agents, concerns, and theory. Evaluates at three levels: isolated prompts, assembled context, and ov…
Testing
shadow-mode-runner
Coordinates SHADOW mode operation where the agent runs parallel to human input without delivering output or incurring billing. Measures agreement rates and generates promotion repo…
DevOps
tool-procurement-eval
Evaluates new tools before adoption by framing needs first, designing trials with upfront success criteria, auditing stack fit and overlap, and conducting security/data reviews sca…
General
new-project-gate
Pre-build gate that evaluates new ideas against both agent criteria (legible, drivable, automatable) and human criteria (usefulness, maintainability) before creating any tool or re…
AI / ML
agento11y-prod-setup
Sets up production evaluation and guardrails for a deployed AI agent using live traffic analysis. Reads agent source and samples real usage to identify missing online evaluators an…
AI / ML
annotate-spans
Records structured annotations on agent spans and traces, coaches review practices, and supports failure taxonomies or human/LLM evaluation workflows. Load when preparing to save f…
Design
ai-product-design
Designs AI features with clear capability boundaries, user controls, feedback loops, recovery paths, and evaluation criteria. Produces task/risk models, interaction architectures, …
Research
mlsys-experiments
Designs and audits MLSys paper evaluations: picks representative workloads/hardware, tunes baselines symmetrically, reports throughput/latency/memory/cost/quality together, structu…
General
capstone-final-eval
Evaluates Ewha Womans University capstone design final deliverables including team presentations, PDF reports, GitHub repositories, and Project Briefs. Use for end-of-term capstone…
Research
acmmm-artifact-evaluation
Prepares ACM Multimedia submissions by packaging code, models, datasets, or media for review or public release, selecting the appropriate track (OSS Competition, Dataset, Reproduci…
Business
recommendation-brief
Generates and audits the final purchase brief: ranked recommendation with confidence rationale, trade-off map, evaluation matrix, sensitivity checks, forensics flags, merchant veri…
Research
light-idea-critique
Evaluates research ideas against top-tier publication standards to distinguish genuine breakthroughs from incremental or superficial combinations. Delivers blind-then-explicit scor…
Automation
loop-design-check
Designs and reviews goal-oriented agent loops for failure modes like token waste, verifier gaming, and completing wrong answers. Writes loops with decidable goals and skeletons; au…
Research
nsdi-experiments
Audits or designs evaluation plans for networked-systems papers: selects traces, testbeds, and deployment evidence, sizes scale and failure-injection experiments, chooses adversari…
Testing
aas-experiments
Builds rigorous experimental evidence for AAS submissions: designs RQ-driven studies, selects fair baselines, analyzes stability/robustness, applies statistical tests, structures a…
AI / ML
calibrator
Calibrates prompts, agents, and skills by comparing AI outputs against author-edited versions. Identifies recurring divergence patterns and generates targeted instruction fixes in …
Research
elicit
Applies one of nine structured reasoning methods—pre-mortem, inversion, first principles, red team/blue team, Socratic, constraint removal, stakeholder mapping, analogical, or seco…
AI / ML
evaluating-cosmos-policy
Evaluates NVIDIA Cosmos Policy on LIBERO and RoboCasa simulation environments. Use when setting up cosmos-policy for robot manipulation evaluation, running headless GPU evaluations…
AI / ML
model-profiler
Profiles models from any provider family against a fixed set of 10 work categories, mapping public benchmarks into tier rankings with full provenance. Emits routing-table.json, aud…
AI / ML
evaluate-agents
Separate runtime outcome, metrics, business evaluation, and release readiness while preventing failures from masquerading as passes. Use when creating, changing, reviewing, or diag…
Research
legalsearchqa-eval
Benchmarks retrieval of current legal information from external sources and reasoning over it for multiple-choice legal questions. Measures factual accuracy, uncertainty calibratio…
Testing
eval-case-author
Generates evaluation cases for any artifact using real human/agent pairs or declared synthetic cases, then persists them in evals/{artifact_id}/cases/case-{n}.md. Supports C4 (≥30 …
Productivity
jobfit
Evaluates job fit by discovering matching roles across boards and scoring provided postings A–F; researches compensation, company signals, and legitimacy, then outputs ranked brief…
Research
devpsych-review-process
Examines Developmental Psychology (APA) manuscript evaluation criteria including masked review, developmental significance weighting, design rigor, JARS reporting, and TOP transpar…
AI / ML
embodied-eval-automation
Automates reproducible batch episode collection, validation, and auditing for policy, VLA, and world-model benchmarks across local, SSH, or cloud GPU hosts. Handles asset reuse, fo…
AI / ML
eval-harness-first
Builds the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when launching fine-tuning,…
Research
raven-research-looped-reasoning-eval
Scaffolds reproducible experiments comparing looped/recurrent-depth transformers against matched vanilla models on vulnerability discovery after security-corpus pretraining. Use fo…
Research
sigcomm-experiments
Designs and audits network systems experiments following ACM SIGCOMM standards: progressing from microbenchmarks to testbed and trace-driven evaluation, selecting strong baselines,…
Research
eccv-experiments
Designs and audits ECCV experimental programs: benchmark selection that holds through September conferences, fair baseline comparisons in the foundation-model era, mechanism-isolat…
Business
vetting-services
Evaluates appraisers, graders, dealers, and auction houses for compliance and credibility. Use when choosing valuation professionals, assessing grading services, reviewing dealer c…
Research
fast-experiments
Designs and audits storage evaluations for USENIX FAST, covering real devices, firmware, aging/preconditioning, standard workloads (SNIA IOTTA, YCSB, filebench, fio), write amplifi…
Research
lancet-fit
Pre-evaluate any study for The Lancet by checking clinical or public-health significance, global reach, practice-changing potential, and equity considerations to determine fit for …
AI / ML
model-checkpoint-evaluator
Evaluates multiple model checkpoints across benchmarks using accuracy, F1, or custom scoring to identify the best performer. Scans checkpoint directories, executes evaluations, com…
Showing the top 60 of 1,282. See the full list →