Designing Better Evals
Build robust evaluation frameworks, benchmarks, graders, and regression tests that meaningfully track model improvement.
First Annual Summit on Model Training and Evaluation
Our speakers have trained and evaluated models at the following labs and companies:
The rapid improvement of frontier models has created a new set of questions around how models are trained, evaluated, and improved after pretraining. These questions increasingly sit at the center of model development, but they remain distributed across research communities, industry teams, and infrastructure providers.
FAIR Summit is a gathering focused on model training and evaluation. It brings together researchers, engineers, and practitioners working across reinforcement learning, human and synthetic data, evaluations, alignment, model behavior, and infrastructure.
Our goal is to create a community for the people responsible for improving models to be more trustworthy, reliable, and capable.
A network of researchers, operators and founders advancing the frontier of AI through post-training and model evaluation.
Build robust evaluation frameworks, benchmarks, graders, and regression tests that meaningfully track model improvement.
Explore judge calibration, bias, agreement, failure modes, and when human evaluation is still required.
Design better rubrics, train evaluators, measure agreement, improve annotation quality, and build reliable adjudication systems.
Understand how supervised fine-tuning, preference optimization, reinforcement learning, and data selection fit together.
Design reward signals, build reward models, prevent reward hacking, and identify where verifiable outcomes can improve training.
Generate, curate, filter, and combine synthetic and human data to create higher-quality training mixtures.
Evaluate agents across long-horizon tasks, tool calling, browsing, coding, research, and real-world environments.
Design environments, tasks, and feedback loops that produce useful learning signals for increasingly capable agents.
Identify systematic failures, build error taxonomies, evaluate robustness, and understand capability tradeoffs.
Determine whether a new model or checkpoint is actually better, and detect when improving one capability causes another to degrade.
Evaluate jailbreaks, refusals, over-refusal, misuse, robustness, and other behaviors critical to trustworthy deployment.
Founder & CEO
Founder & CEO
PerleGlobal Head of Adoption
Co-Founder
MidcenturyGTM | Account Management
Forward Deployed TPM
MicrosoftHead of API Products
VEEDRevealed in waves · 2027
Revealed in waves · 2027
Revealed in waves · 2027
If you're doing frontier work in model research, training, or evaluation, apply to get involved.
Support the researchers shaping how frontier models are trained and evaluated.
Summit hosts are announced in 2027. AI Circle’s current partners include Prolific, Invisible and Encord.