← All tasks
Data processingPay for results

Evaluate a sample of personalized AI responses in English

Project brief

We need an English-speaking AI evaluator for a small, self-contained personalization test. You will create 30 realistic one-to-five-turn prompts from supplied fictional user profiles, compare two model responses for each prompt, and assess relevance, accuracy, helpfulness and appropriate use of personal context.

Flag hallucinations, unsupported assumptions and forced personalization. For every comparison, select the stronger response and write a concise rationale that points to the relevant turn. Follow the supplied rubric and data-hygiene instructions throughout; no personal account data is required for this commission.

Deliverables & acceptance

What you'll deliver

  • Completed evaluation file containing 30 prompt-and-response comparisons
  • Preference decision and concise rationale for every comparison
  • Issue annotations covering hallucinations, unsupported assumptions and forced personalization

What the result must meet

  • Thirty prompt-and-response comparisons are completed against the supplied fictional profiles and evaluation rubric.
  • Every comparison includes ratings for relevance, accuracy, helpfulness and personalization, plus a clear preference rationale.
  • Hallucinations, unsupported assumptions and forced personalization are explicitly identified, and the completed file passes the supplied completeness checks.
Browse more tasks →How pay-for-results hiring works →

Download RenX

Get the app.

Install, sign up, and start on your free plan with welcome credit included. No credit card required.

On a platform not listed? Leave your email and we'll notify you when a build is available.