eval-harness-first
자동 수집 항목입니다. 설치 설정은 저장소 README에서 그대로 가져왔지만, 한국어 설명과 부서 분류는 아직 확인하지 않았습니다. 설치 전에 저장소를 한 번 확인하세요.
모든 파인튜닝 실행을 검증하는 평가 하네스를 구축합니다 — 골든 세트, 실패 유형별 채점기, 심사자 보정, 베이스 모델 기준선을 포함합니다. 파인튜닝 작업을 시작할 때, 트레이스를 평가 세트로 변환할 때, 또는 심사자를 사람 라벨과 대조하여 보정할 때 사용합니다.
설치하기
Claude Code에서 위 두 줄을 차례로 실행합니다. `llm-finetuning` 묶음에 이 스킬이 들어 있습니다. 설치한 뒤에는 이름을 언급하면 발동합니다 — 예: "eval-harness-first 스킬로 정리해줘". 원본: https://github.com/wshobson/agents/tree/main/plugins/llm-finetuning/skills/eval-harness-first
/plugin marketplace add wshobson/agents /plugin install llm-finetuning@claude-code-workflows
설치와 실행은 사용자 환경에서 이뤄집니다. team-ai는 이 도구를 호스팅하거나 대신 실행하지 않습니다. 제공처와 저장소를 확인한 뒤 설치하세요.
설명
원문: Build the evaluation harness that gates every fine-tuning run — golden sets, per-failure-mode graders, judge calibration, and base-model baselines. Use when starting a fine-tuning effort, when converting traces into an eval set, or when calibrating a judge against human labels.