Hundreds of contractors are reading real ChatGPT conversations — including, sometimes, deeply personal ones — and rating the chatbot’s responses to make its next version better behaved. According to a 404 Media investigation published this week, the work runs under the codename Project Lily, and the people doing it are paid more than US$50 an hour.
Strip away the privacy debate and the report contains a genuine jobs story: one of the most consequential tech companies of the decade is building a new profession from scratch, and it is hiring ordinary people — not just PhDs — to do work a model cannot do to itself.
What the job actually is
The contractors follow a three-stage routine, according to instruction guides seen by 404 Media. They read a real user’s prompt, summarise what that person was actually trying to achieve, then score four ChatGPT responses on a one-to-seven scale — from “unacceptable, unusable” to “would be hard to meaningfully improve” — with a written rationale.
The guides read like a style manual for artificial manners. Reviewers are told to penalise “AI-speak” and emoji misuse, to flag sycophancy, engagement-bait endings and “patronising assumptions,” and to reward responses that are “smart, but humble.” One document instructs that the model should match the user’s tone “but slightly less intensely.”
It is judgment work, performed at industrial volume. OpenAI told 404 Media it runs conversations through a “Privacy Filter” model before reviewers see them, while acknowledging that filter “can make mistakes” — a caveat that fuels the story’s privacy angle, and the reason the company updated its help pages after being contacted. Users can turn off the default-on “improve the model for everyone” setting, though it applies only to new conversations.
The profession behind the codename
The numbers suggest this is bigger than one company’s evaluation pipeline. Deel’s analysis of more than a million worker contracts across 37,000 companies found general AI trainer roles grew 283 percent cross-border in 2025. The occupation now spans more than 70,000 workers at over 600 organisations — from basic annotators labelling images to subject-matter experts reviewing a model’s medical or legal reasoning.
Pay in the field spans an enormous range. The same hour can be worth roughly US$4 for entry-level labelling or US$200 for formal proof-writing, and Deel’s data puts typical AI trainer pay between US$15 and US$75 an hour. What separates the tiers is not seniority but expertise: whether you can defend a judgement call in a domain a model cannot grade from a rubric alone.
Anthropic confirmed to 404 Media that it uses human reviewers to improve Claude’s responses, and Google’s Gemini carries a disclaimer that “humans review some saved chats.” The work is now a standard part of how frontier models get made — and a career category that barely existed three years ago.
New Zealanders are in the hiring pool
This is one corner of the AI economy where geography barely matters, and NZ workers are actively being courted. Global AI-training recruiter Mercor lists remote roles open to New Zealand applicants, including a Legal Expert – AI Trainer posting advertising US$90–130 an hour — roughly NZ$150–215 at current exchange rates — for lawyers who can critique a model’s legal reasoning. Local job boards show similar demand: SEEK currently lists 548 roles under its “AI trainer” search in New Zealand, though that catch-all includes conventional staff trainers alongside genuine AI-evaluation work.
The pattern matches what Deel found globally: the work recruits through subject expertise, not tech credentials. A nursing qualification, courtroom experience, or genuinely native command of a language all qualify. Knowing how to prompt well does not.
There is a two-sided lesson in the data. Meta’s Dublin AI-training operation cut around 700 contractor roles in April as internal systems replaced the very people who trained them — evidence that parts of this profession can be automated out from under itself. And UCLA professor Sarah T. Roberts, who studies hidden digital labour, told 404 Media the generous pay is “for now,” arguing the industry treats the human judgement it depends on as its least valued input. Anyone treating AI training as a durable career should read both stories and decide for themselves where the durable part is.
Who should look at this, honestly
The realistic read: this is a strong side-income or bridge role with a real skills ladder, not a guaranteed career. The volume-annotation tier is exposed to exactly the automation it feeds. The expert-evaluation tier — where lawyers, clinicians, engineers and native-language speakers review specialist output — is where the money concentrates, and that work rewards knowledge built outside the AI industry. For professionals whose expertise is portable, it is a new, flexible export market for judgement. For those without a defensible domain, the entry tier is a bridge, not a destination.
There is also an unavoidable caveat for any prospective applicant: the work means reading other people’s private conversations at scale. Some will find that disqualifying on principle; regulators in multiple jurisdictions are already scrutinising it. That discomfort is part of the job description.
The bigger signal for New Zealand workers is what the job category proves: the AI boom is hiring humans for precisely the things models can’t self-assess — taste, tone, domain judgement, cultural nuance. The Economist’s recent analysis put AI-created jobs in America around the one million mark, with LinkedIn pointing to roughly 640,000 new AI-specific roles since 2023. Prompt review is a small, unglamorous slice of that count — but it is a slice a New Zealander with the right expertise can do from a kitchen table in Tauranga.
FAQ
How much do AI trainers get paid in 2026? Rates range from about US$4 an hour for basic image labelling to US$200+ for expert-level work. Deel’s contract data shows typical AI trainer pay between US$15 and US$75 an hour, while OpenAI’s Project Lily reviewers were paid more than US$50 an hour. In New Zealand dollars, specialist listings run as high as NZ$150–215 an hour.
Do you need a tech background to become an AI trainer? No. Recruiters and labs consistently say subject-matter expertise — medicine, law, engineering, translation, finance — is what commands premium rates. General prompting skill is assumed, not paid for.
Is AI training a real job or a gig? It is both. Deel counts 70,000+ workers across 600+ organisations, but most engagements are contracts, and lower-tier annotation work can be automated. Treat it as flexible income with a skills ladder, not a lifetime role.
Can New Zealanders apply for AI trainer roles at OpenAI or other labs? Directly with the labs, usually not — the work flows through intermediaries like Mercor and Crossing Hurdles, which do list roles open to New Zealand-based contractors. Pay and eligibility vary by project, and applicants should read each platform’s terms carefully.
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.