Ryan Marten builds Harbor, a framework for evaluating agents and running RL environments, and Terminal-Bench, a benchmark for AI agents on terminal tasks, where he is part of the project leadership.
Ryan Marten builds Harbor, a framework for evaluating agents and running RL environments, and Terminal-Bench, a benchmark for AI agents on terminal tasks, where he is part of the project leadership.