Skip to content

[feat] Evaluate a hosted LLM API as an alternative to on-LAN Ollama for the judge #23

Description

@dogkeeper886

Context

The LLM judge currently targets a local Ollama instance (`gemma3:4b` via `LLM_JUDGE_URL`). FR-005 listed Moving to a hosted LLM API as a Non-goal because it changes the judge's cost / latency / governance profile. This ticket picks that evaluation up as a standalone decision.

Why consider it

  • Removes the self-hosted Ollama infra burden (no runner, no GPU box, no model disk space, no patching).
  • Access to stronger models — frontier judges (Claude, GPT-4o) would likely catch semantic regressions that `gemma3:4b` misses, closing FR-005's documented "narrow risk class" gap.
  • No LAN requirement → CI can run on GitHub-hosted runners end-to-end, dropping the self-hosted runner path entirely.

Why not

  • Per-test cost. Full suite has 19+ tests; if we run the pipeline on every PR, the monthly bill is real and load-bearing for any decision to re-enable automatic triggers.
  • External dependency + potential data leakage. The judge sees log fragments, step output, and test context. Need to decide whether that's acceptable to send to a third-party provider.
  • Latency is worse than local Ollama for short prompts (round-trip to hosted API vs localhost).

Decisions needed (discuss in issue comments, then capture in an FR if we proceed)

  1. Pick a provider. Claude / OpenAI / other. Cost per judgment, model quality, data retention policy.
  2. Is the test data safe to send? Log fragments include API keys masked, XML-RPC responses, TestLink entity names. Nothing production-sensitive today, but the policy should be explicit.
  3. Cost ceiling. What's the monthly budget? Drives whether this is for PR CI or just manual dispatches.
  4. Keep Ollama as an option? Env-var switched, or ripped out?

Must (once a decision lands)

  • An FR-NNN document capturing the provider choice, data-handling policy, and cost analysis.
  • One provider wired behind the existing `LLMJudge` interface (same `judge(test, stepOutputs)` contract).
  • Selectable at runtime: env var or CLI flag to pick provider.
  • `CLAUDE.md` + `CI_SETUP.md` updated with the provider secret names.

Out of scope

  • Multi-provider abstraction or provider-switching logic beyond a basic selector. Pick one to land first.
  • Self-hosted runner work (separate ticket).

Acceptance criteria

  • Decision captured in an FR, linked from this issue's Spec: line.
  • If we proceed: CI can dispatch `test-pipeline.yml` with the hosted judge and complete end-to-end, using only GitHub-hosted runners.

Blocks / unblocks

  • Alternative path to the self-hosted runner ticket. If this one lands, the self-hosted one may become unnecessary.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions