콘텐츠로 이동

arXiv 2604.16790 — Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering

제목: Bias in the Loop: Auditing LLM-as-a-Judge for Software Engineering 베뉴: arXiv (2026) 한 줄: judge가 불확실할수록 톤·권위 cue에 앵커링된다 — confidence/자신감 신호를 judge에 노출하면 판단이 +30%p까지 흔들린다는 실증. ⚠️ 도메인: 코드 평가. fact verification 일반화 시 도메인 차이 감안 필요.


다루는 12개 Bias

군 bias
신뢰성/권위 Authority, Model-name, Refined
톤/확신성 Sentiment(자신감 톤), CoT, Verbosity
위치/제시 Position, Bandwagon, Distraction, Self-enhance, Final-only, Diversity

핵심 실증 — 자신감 톤이 판단을 뒤집는다

Sentiment(자신감 있는 톤) 주입 효과 (Qwen2.5-Coder-3B):

작업 기준 일관성(CR) Sentiment 주입 후 변화
CodeGen 58.57% 86.12% +27.55%p
CodeRepair 65.75% 87.08% +21.33%p
TestGen 50.36% 80.56% +30.2%p

"When the judge is inherently uncertain, strong linguistic or tonal cues act as an anchor."

메커니즘: judge가 애매할 때(기준 CR ≈ 50%), 자신감 톤이 실제 코드 품질 평가를 압도한다. "일관성은 높지만 체계적으로 편향된" 결정 → 위험한 신뢰성 환상.

"Authority and bandwagon biases cause judges to favor responses with citations, a confident tone... even when those signals are irrelevant or fabricated."


불확실성 × bias 증폭

  • judge 기준선이 불확실할수록(CR 50% 근처) 표면 cue 앵커링 극대화.
  • Hard 난이도에서 Sentiment/CoT/Refined의 A-B 정확도 격차 60–85%p.

Mitigation — judge 독립성 확보

  1. Position swap: 모든 평가를 orig/swap 두 번 → 위치로 결과 바뀌면 편향 신호.
  2. 단일 bias 주입 통제 측정: baseline 대비 편차 추적.
  3. 신뢰도 임계 기반 에스컬레이션: CR < 70% & Hard → 인간 검증. TestGen CR < 60% → 자동 judge 단독 금지.
  4. judge 체인: 단독 judge → 불일치 시 컴파일·정적분석·테스트 등 외부 검증.

"reported judge performance may reflect prompt artifacts rather than stable assessment ability"


본 프로젝트와의 연결

  • ④ "judge에 confidence/logprob 주입 = 안티패턴"의 강력한 근거: 자신감 톤 노출만으로 judge가 +30%p 흔들림 = confidence를 judge 입력에 넣으면 검증이 무뎌진다는 실증.
  • mitigation 일치: "judge는 외부 근거만 보고 독립 판단" = 톤 무시 지시 + isolation + swap.
  • 시사: logprob은 judge가 보는 입력이 아니라 시스템이 검증을 배치하는 메타 정보로만 써야 한다 → "스케줄러지 참고인 아니다" 원칙의 근거.
  • ⚠️ 코드 평가 도메인이므로 fact-claim 검증으로 옮길 때 효과 크기는 재측정 필요.