Machine Learning Engineer - Model Evaluation & Experimentation
weekday-1 · United States
About The Role
# Machine Learning Engineer - Model Evaluation & Experimentation
> Weekday AI · United States (Remote) · Part-time · Posted 2026-07-30
**Workplace:** remote
**Department:** AI Training
## Description
**This role is for one of our clients**
**Compensation: $60-$90 per hour**
Join a pioneering AI initiative focused on building the next generation of evaluation benchmarks for frontier AI models. We are seeking experienced Machine Learning Engineers and Researchers to bring hands-on expertise in model development, experimentation, and evaluation to create rigorous benchmark tasks for advanced AI systems.
In this role, you will design sophisticated, multi-step machine learning challenges inspired by real-world research workflows. From implementing experimental ideas and running training pipelines to analyzing model behavior and validating results, you will help establish high-quality evaluation benchmarks that reveal the strengths and limitations of frontier AI models.
This is a **fully remote, full-time engagement** requiring approximately **35 hours per week**.
## Requirements
### Key Responsibilities
- Design realistic machine learning benchmark tasks based on research workflows, including model implementation, experimentation, training, evaluation, and performance analysis.
- Translate open-ended research concepts into structured, reproducible evaluation tasks with clearly defined success criteria.
- Implement machine learning solutions using **Python**, execute experiments, and produce reference implementations that demonstrate correct methodology and expected outcomes.
- Develop benchmark tasks involving reinforcement learning concepts such as reward functions, policy optimization, training dynamics, and model behavior where applicable.
- Evaluate AI-generated solutions by identifying implementation errors, experimental flaws, incorrect reasoning, and unsupported conclusions.
- Collaborate with AI researchers and fellow subject matter experts to continuously improve benchmark quality, technical rigor, and evaluation consistency.
### Required Qualifications
- Master's degree, PhD, or equivalent practical experience in **Machine Learning, Computer Science, Artificial Intelligence, Data Science, or another quantitative STEM discipline**.
- Minimum **1 year of professional experience** in machine learning research, research engineering, applied AI, or another research-intensive technical role.
- Strong hands-on experience designing, training, evaluating, and optimizing machine learning models through complete experimental workflows.
- Practical experience conducting machine learning experiments, including experiment setup, hyperparameter tuning, execution, validation, and analysis.
- Strong understanding of modern **Large Language Models (LLMs)**, their capabilities, limitations, and evaluation methodologies.
- Proficiency in **Python** and **Git**, with experience working in both script-based and notebook-based development environments.
- Familiarity with reinforcement learning concepts—including reward functions, policy optimization, and training behavior—is preferred.
- Experience with AI evaluation, benchmark development, AI training, or task authoring is highly desirable.
- Excellent analytical thinking, creativity, attention to detail, and the ability to solve complex, open-ended technical problems independently.
- Strong written communication skills for documenting experimental methodologies and technical findings.
- Ability to commit approximately **35 hours per week** on a consistent basis.
### Preferred Qualifications
- Experience developing or evaluating large language models, foundation models, or generative AI systems.
- Background in reinforcement learning, deep learning, distributed training, or model optimization.
- Familiarity with benchmark design, AI safety evaluations, or research-quality experimentation.
- Experience contributing to research publications, open-source machine learning projects, or advanced AI systems.
### Why Join
- Help shape how next-generation AI systems are evaluated through rigorous machine learning experimentation.
- Collaborate with leading AI researchers developing frontier evaluation benchmarks.
- Apply your expertise to improve AI reasoning, model quality, and experimental reliability.
- Contribute directly to benchmark development that advances the capabilities of state-of-the-art AI systems.
- Enjoy the flexibility of a fully remote engagement while working on impactful AI research initiatives.
### Equal Opportunity
We are committed to fostering an inclusive and diverse environment where all qualified applicants receive equal consideration. Reasonable accommodations are available throughout the application and engagement process.
### Contract & Engagement Details
- Independent contractor engagement.
- Fully remote with flexible working hours.
- Expected commitment of approximately **35 hours per week**.
- Project duration may be extended, shortened, or concluded based on project requirements and individual performance.
- Work does not require access to confidential or proprietary information from any current or former employer.
- Payments are issued weekly based on approved work completed.
- At this time, we are unable to support H1-B or STEM OPT candidates.
## Apply
[Apply at Weekday AI](https://apply.workable.com/weekday-1/j/BF32A485B8/apply)
---
Powered by [Workable](https://www.workable.com)
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring