Senior RL Engineer Needed
Summary
I'm competing in the Humanoid Olympics, a simulated athletics competition, and I'm looking for an experienced reinforcement learning freelancer to build a winning policy.
The task is to train one policy that controls a Unitree G1 humanoid (12 actuated leg joints, no arms) across six events in MuJoCo:
- 100 m sprint
- 400 m circular sprint
- 100 m hurdles (barriers up to 1.15 m)
- High jump (bars up to 1.30 m)
- Long jump (6 m void)
- Triple jump
Surface friction (μ 0.30–1.25) and wind (up to 8 m/s) change every round and are not observable, so the policy must adapt through its 256-float recurrent state. Fouls, falls, and leaving the lane score zero or near zero.
The full environment, referee, and baseline policy will be provided.
The goal is to reach the top of the leaderboard.
Scope of Work
- Set up the official environment and reproduce the baseline.
- Train a robust policy using domain randomization, recurrent memory, and curricula (teacher-student or privileged training is welcome).
- Build a multi-seed evaluation script that reports per-event and per-condition scores and failure reasons.
- Export the final policy to ONNX with the exact signature.
- Make sure the model is under 15 MB and finishes a full 24-attempt meet on 2 CPUs within 900 seconds.
- Deliver the training code, config, checkpoints, and a short README so I can retrain and resubmit.
Required Skills
- Proven experience training legged or humanoid locomotion policies with RL (PPO or similar)
MuJoCo, Isaac Lab, or Brax
- Domain randomization, reward shaping, curriculum learning
- Recurrent policies (GRU/LSTM) for unobserved dynamics
- PyTorch and ONNX export
Nice to have: Unitree robots, jumping or parkour skills, CPU inference optimization, Bittensor experience.
Terms
- All code, models, checkpoints, and training artifacts created for this project are owned by the client.
- The freelancer agrees not to submit this policy, or any derivative of it, to the competition from their own or any -third-party account.
- Work and results are confidential until the client decides otherwise.
To Apply, Please Include
- Links or videos of locomotion policies you have trained (required).
- A short plan (5–10 sentences): how would you train one policy to both sprint and jump when friction is hidden and changes every attempt?
- Your estimated timeline and whether you prefer fixed-price milestones or hourly.
Proposals without examples of previous locomotion work will not be considered.