Posted last month
Research Engineer on the Code RL team at Anthropic, focusing on reinforcement learning to improve code generation and model reasoning. Work involves designing RL environments, conducting experiments, and scaling RL infrastructure.