Source description
About the role
About HUD
HUD https://www.hud.ai/ is building infrastructure to create RL training data and evals for frontier AI agents, as well as a marketplace to sell these to frontier labs through the HUD marketplace. Our platform is used by frontier labs, Fortune 500 companies, and startups. We’ve raised $16M from top VCs and were YC W25.
About THE Role
We’re looking for Research Engineers to build high-quality benchmarks for evaluating frontier agents on domain-specific tasks. You’ll build benchmarks that are technically rigorous, practically useful, and credible to frontier labs.
Responsibilities
-
Own the design, implementation, and quality of HUD’s internal agent benchmarks
-
Work with subject-matter experts to define tasks and create domain-specific benchmarks that evaluate agents on realistic workflows
-
Build infrastructure to reliably run models and agents against benchmark tasks
-
Develop metrics and analyses to understand benchmark difficulty, reliability, and failure modes
-
Validate whether benchmark performance correlates with real-world evals, customer needs, and lab expectations
-
Write clear documentation and benchmark reports that make results legible and credible to technical audiences
-
EXPERIENCE
-
You may be a good fit if you have:
-
Proficiency in Python, Docker, and Linux environments
-
Published papers or written technical blogs on relevant topics such as public benchmarks and their limitations, model failure modes, etc. - please link in your application
-
Strong understanding of what a “good benchmark” means and what makes one realistic, reliable, and useful
-
Experience working on environments and evals
-
Curiosity and ability to truly understand how workflows in various domains work
-
Strong candidates may also:
-
Be detail-oriented and able to spot subtle inconsistencies or edge cases in tasks
-
Be able to reason from first principles about task design, scoring, and failure modes
-
Thrive in unstructured problem spaces
-
Early-stage startup experience with ability to work independently in fast-paced environments
-
Strong communication skills for remote collaboration across time zones
-
We prioritize technical aptitude and learning potential over years of experience. Motivated candidates are encouraged to apply even if they don't meet all criteria.
-
TEAM & COMPANY DETAILS
-
Team Size: ~15 people currently, mostly full-time in-person, but some remote.
-
Our team: Our team includes 4 International Olympiad medalists (IOI, ILO, IPhO), serial AI startup founders, and researchers with publications at ICLR, NeurIPS, etc.
-
Company stage: We have 8 figures in funding and high revenue growth. We’re scaling profitably and quickly to meet very strong demand.
-
LOGISTICS
-
Employment: Full-time.
-
Location: On-site only, for now. You can join the team in the San Francisco Bay Area or Singapore offices.
-
Visa Sponsorship: We provide support for relocation and visas for strong full-time candidates to the US or Singapore.
-
Timeline: Applications are rolling. The process is 2 technical interviews and a 2-3 day work trial.
What WE Offer
-
Competitive compensation
-
100% covered top-of-the-line medical, dental, and vision from Blue Shield of CA (US employees)
-
Lunch and dinner when you’re in the office
-
Company-wide holiday break (Christmas Eve to New Year’s Day) on top of PTO and paid holidays
-
Other perks including an Equinox membership, 401k, and commuter benefits (US employees)
-
Unlimited* access to tokens for ChatGPT, Claude Code, Cursor, etc. *By unlimited, we mean no one on our token usage leaderboard has ever hit a limit. So we have no idea what the limit is.
-
Due to high volume, we may not actively respond to every application, but feel free to contact us at recruiting@hud.so or elsewhere if we missed your application!
More at hud
Related open roles
Research Engineer, Benchmarks
Remote · San Francisco Bay Area · Singapore
Applied Research Engineer
Remote · San Francisco Bay Area · Singapore
Research Engineer (General)
Remote · San Francisco Bay Area · Singapore
Research Engineer, Synthetic Data
Remote · San Francisco Bay Area · Singapore
Research Engineer, Synthetic Data
San Francisco Bay Area · Singapore · Onsite
Research Engineer, QC Automation
San Francisco Bay Area · Singapore · Onsite
