Build a Python tool to compare repeated LLM runs
Project brief
Create a lightweight Python prototype that repeats one LLM task 5-10 times, stores outputs and available state, and identifies differences. Apply an agreed rule set to label runs as stable, boundary or violation. Keep it a small script or app, not a full platform.
Delivery deadline: 2026-10-05. Expected production time: 7 calendar days after contract funding and receipt of agreed materials. Confirm the scope and materials before accepting the contract.
Project materials will be shared with shortlisted applicants during scoping and supplied before work begins.
Required project materials: Authorized API account, task prompt, classification rules and example fixtures.
Submit the work and acceptance evidence through RenX. Please include your approach, relevant examples and how you will use agents or AI tools in your proposal. Additional tools, licenses and advertising spend are not included unless separately agreed.
Deliverables & acceptance
What you'll deliver
- Runnable Python tool, JSON run logs, comparison output and setup instructions.
What the result must meet
- One command runs the agreed repetitions and saves a record for each run.
- Supplied test fixtures produce the expected difference labels; API keys stay outside committed code.