We need native Afrikaans speakers to transcribe conversational audio recordings for an AI training dataset. Transcripts must be fully human-generated or human QA-ed — no machine-only transcripts accepted.
What you'll do:
• Transcribe conversational Afrikaans audio files (2 minutes to 1 hour per file).
• Deliver transcripts in JSON format with:
→ Time-coding (no overlaps across speakers)
→ Speaker diarization (each speaker individually labeled, tied to Speaker IDs)
→ Sequential ordering of utterances
→ Minimum 95% precision
• Confirm each transcript is fully human-generated or human QA-ed.
Workflow:
• We provide the audio file plus a machine-generated draft transcript (from our Amazon Transcribe pre-pass) where the ASR handles the language reasonably. For languages where ASR performance is poor, you'll transcribe from scratch.
• You listen to the full audio, correct every error in the draft (or transcribe from scratch), verify speaker labels and timestamps, and output the corrected JSON.
• You attach an attestation confirming the transcript is human-generated or human QA-ed.
Requirements:
• Native or near-native speaker of Afrikaans.
• Comfortable working with JSON output (schema and conventions provided in our Transcriber Guide).
• Access to a laptop or desktop computer.
• Ability to deliver on rolling schedule aligned with 20-day tranches.
What you'll be paid:
• $5 per accepted audio hour transcribed, paid per accepted submission.
• Accepted = passes our internal QA sample check (we spot-check a portion of your files).
Timeline: Rolling submissions from now through Sep 14, 2026. Tranche 1 due August 4th.
What to submit:
• JSON transcript file per audio file (schema in the Transcriber Guide we'll share on acceptance).
• Attestation that the transcript is human-generated or human QA-ed.