AI evaluation · Research platform
Code Grading Evaluation Platform
A repeatable Python workflow for comparing OpenAI and Gemini grading decisions against rubric expectations and human review. The system flags inconsistent outputs and supports threshold-based, human-in-the-loop evaluation.