LLM Evaluation post training

via Freelancer ·

Budget / SalaryHourly project
TypeFreelance project
LocationRemote
Posted1 hour ago
Senior ML Engineer / Advisor / Technical co-founder: Model Post-Training & Alignment

About Bentham Research
Grounding Machine Intelligence in the Humanities

Bentham is an applied research lab working at the intersection of AI and the humanities. We build doctorate-authored, peer-reviewed evaluation instruments, training environments, and datasets for frontier AI labs, enterprises, and governments. Our core focus spans ethics, moral reasoning, philosophy, political theory, law, theology, and history. We believe expanding model capabilities requires anchoring machine judgment in human wisdom through bottom-up, scholar-led workflows paired with AI-in-the-loop validation.

The Role
We are seeking an ML Research Engineer or Technical Advisor with hands-on experience in post-training models at a major frontier AI lab. You will bridge our team of PhD humanities scholars and technical alignment workflows, moving Bentham from core methodology pressure-testing into active project execution. You will help design, build, and validate the pipelines that translate complex humanities rubrics into high-yield evaluation benchmarks and post-training datasets.

Key Responsibilities

Technical Architecture: Translate scholar-authored, rubric-based datasets into machine-readable formats optimized for SFT, RLHF/RLAIF, DPO, and reward modeling.
Methodology Pressure-Testing: Scrutinize and refine our evaluation framework to ensure our scholar-led datasets stand up to the technical standards of frontier lab eval teams.
Pipeline & Environment Development: Lead the development of pilot training environments and benchmarking tools that test model capabilities beyond traditional STEM domains.
Cross-Domain Collaboration: Interface directly with Bentham CEO Marcus Heal and doctorate domain experts to translate qualitative human reasoning into rigorous, verifiable alignment signals.

Qualifications

Prior experience in model post-training, preference tuning, or evaluation design at a major AI lab (e.g., OpenAI, Anthropic, Google DeepMind, Meta etc).
Deep technical understanding of SFT, RLHF/RLAIF, LLM-as-a-judge evaluation frameworks, and psychometric benchmark design.
Ability to translate nuanced qualitative criteria (law, philosophy, history) into precise ML feedback loops.
Pragmatic, developer-first mindset with experience taking experimental evaluation hypotheses into production-ready pipelines.
machine learning (ml) ai model development ai research large language models (llms) llm fine-tuning
Apply on Freelancer →

Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.