Human Company builds expert-grade reinforcement-learning environments used to train and evaluate AI models on inference engineering. We are seeking an ML Systems Engineer to own inference-systems tasks end to end: choose a realistic objective, build the execution environment, implement robust ground-truth verifiers, and perform rigorous QA on fidelity, ambiguity, exploit resistance, and difficulty calibration.
This bounty is a paid, four-hour technical work sample for candidates based in India. Strong submissions may lead to a remote full-time or monthly contract role with a broader compensation range of ₹22-110 LPA, calibrated to experience and scope.
Strong candidates have operated production inference systems and owned latency, throughput, or memory outcomes. Depth in at least one of the following is required, with working knowledge across the rest: serving stack internals (continuous/in-flight batching, chunked prefill, scheduling, admission control, KV-cache management); model efficiency (quantization, speculative decoding); distributed execution (tensor, pipeline, sequence, and expert parallelism, disaggregated prefill/decode, multi-replica balancing, elastic inference); hardware and compilation (GPU/FPGA/ASIC kernels, TensorRT-LLM, TVM, roofline-guided optimization); and statistically sound profiling and benchmarking. Contributions to vLLM, SGLang, TensorRT-LLM, or comparable systems are especially relevant.
The accepted candidate will produce a compact design and verifier prototype for a realistic inference-engineering RL task. Do not share proprietary code or confidential employer information.