tokens&
For enterprises
tokens&

Find tools, check provider offers, save a build plan, and share your work when you’re ready.

For buildersFor enterprises

For builders

  • Startup credits and perks
  • Agent Skills
  • Publish a project

For enterprises

  • Start free company workspace
  • Submit a tool, product, or perk

Community

  • Community
  • Newsletter
  • Events
Xin

© 2026 tokensand, LLC. All rights reserved.

  • Terms
  • Privacy
  • Security
  • Data Processing
  • Status
  1. Hackathon
  2. Project gallery
  3. JudgeAndJury
Anonymous builder2 months agoJudging locked: Self-Evolving Agents Hackathon

JudgeAndJury

An LLM Judge that improves over usage that improves user coding intent by coding agents.

Review the project

Start with the source code, then open the demo or video if available.

View GitHub repository
Visit project website
Project description
I have an LLM Judge that tracks coding agent turns and user input and judges the coding agent output and if it's satisfactory for the user. The goal is to reduce LLM coding friction and increase coding agent understanding of user intention. The platform also supports trace-tracking for when errors occur, or when critical infrastructure problems arise. The traces are logged, and using deterministic checks, it would not occur again. So, in total, it tracks errors, mistakes, coding agent turns, user turns, and improves the LLM judge over time through these traces. Users are also able to fine-tune their own model through these traces and specifically create a model just for themselves.
Tools used
  • Actian logoActian
  • Pioneer logoPioneer
  • Replay.io logo
Watch demo video
Project gallery
Project links
  • GitHub repository
  • Project website
  • Demo video
Tools used
  • Actian logoActian
  • Pioneer logoPioneer
  • Replay.io logoReplay.io
  • SSenso
Replay.io
  • SSenso