← All opportunitiesApply for this role
Research Scientist — Evaluation
$85–$150/hrRemotePart-timeResearch8 openingsPosted this week
About this work
For researchers who want to shape how model capability is measured in their own field, rather than have it measured by generalists.
What you’ll do
- Design evaluation problems at the frontier of your field
- Assess model reasoning for subtle domain errors
- Identify misconceptions a non-specialist reviewer would miss
- Advise on what a meaningful benchmark in your field looks like
What we’re looking for
- A PhD or equivalent research experience in any discipline
- Active familiarity with current work in your field
- No AI background required
Time commitment
5–15 hours per week, projects typically run 4–10 weeks
Pay and how you’re paid
$85–$150 per hour, paid every two weeks by bank transfer
How to apply
Apply once and we’ll consider you for this role and anything else that fits. The full process is set out on How it works.
Similar opportunities
Clinical Medicine Reviewer
NewReview AI-generated clinical reasoning against real standards of care, and write the cases that test it.
MedicineRemotePart-time
$90–$150/hr6 openings · Posted this week
ViewRadiology Report Specialist
Assess how well AI describes and interprets imaging, and where its language would mislead a referring clinician.
MedicineRemotePart-time
$110–$160/hr3 openings · Posted 2 weeks ago
ViewPharmacy & Medicines Reviewer
NewTest AI on dosing, interactions and contraindications — the places where a plausible answer does real harm.
MedicineRemotePart-time
$80–$130/hr4 openings · Posted this week
ViewReady to put your expertise to work?
Apply now
Applications reviewed within 48 hours