The tester withdrew from the QA assignment after starting, and the spot was reopened for a replacement.
AI Benchmark Task Author & Software Engineer | Docker, Terminal Agents, Spring Boot
I build reliable software systems and high-quality evaluation environments for AI agents. My work includes designing containerized, machine-verifiable terminal tasks with deterministic artifact contracts, reference solutions, isolated verifiers, and adversarial checks that prevent shortcut or no-op solutions. I create Dockerized benchmark environments, reproducible test harnesses, and curated training and validation data for terminal-agent evaluation. I also have software engineering experience across Flutter, Spring Boot, PostgreSQL, Kafka, AWS, and real-time systems. I have built cross-platform Android and iOS applications, fault-tolerant real-time duel systems using Centrifugo and PostgreSQL advisory locks, and clickstream/attribution pipelines using Kafka, S3/Parquet, and DuckDB. I can help with AI benchmark design, terminal-task authoring, Docker-based testing, backend engineering, data pipelines, and practical AI integrations.
Frontier Bench Task Author at AfterQuery / Kepler (Jul 2026 - Present) AI Benchmark Task Author at Handshake - Dynamo & Seal Projects (Jul 2026 - Present) Software Engineer Intern at Intellectsoft (Jan 2023 - Apr 2025)
The tester withdrew from the QA assignment after starting, and the spot was reopened for a replacement.
The tester submitted two recordings and a screenshot and documented three useful issues involving job loading, password recovery, and unexpected logout. However, the written report retained placeholder timestamps and left much of the requested coverage undocumented, limiting reproducibility and clarity.