LLM Evaluator & AI Red Teamer | Multilingual AI Safety and Quality Specialist
I evaluate and stress-test large language models and generative AI systems for quality, safety, factual accuracy, reasoning reliability, and instruction adherence. I specialize in multilingual English-Polish AI evaluation, reviewing outputs for linguistic accuracy, cultural relevance, consistency, bias, hallucinations, and unsafe behavior. My work includes manual adversarial testing using prompt injections, jailbreak chains, exploit scenarios, edge cases, and stress tests to uncover model vulnerabilities, policy bypasses, reasoning failures, and behavioral inconsistencies. I provide structured, rubric-based feedback, risk reports, and actionable recommendations to improve model reliability and deployment readiness. I have evaluated conversational and multimodal AI systems including Grok, Gemini, and Midjourney, and work effectively in remote, fast-paced environments that require independent judgment, analytical reasoning, pattern recognition, and precise QA reporting.
AI Evaluator at Contract, Remote (Nov 2025 - Jan 2026) AI Evaluation and Adversarial Testing at Freelance, Remote (Jan 2022 - Present) AI Quality & Evaluation Analyst at Contract, Remote (Jan 2022 - Present)