About this project

A research initiative by the Test Community Network to understand how today's large language models perform on professional multiple-choice item generation.

Research purpose

The leaderboard gathers research data from assessment professionals and subject-matter experts. Its goal is to measure how well existing large language models can create valid, high-quality multiple-choice items in domains that require real-world expertise.

Who we need

We are looking for professionals who write, review, or use assessment content: item writers, psychometricians, examiners, instructional designers, trainers, educators, and domain experts. Your judgments help us build a more meaningful picture of model capability than automated benchmarks alone can provide.

How to contribute

You create a multiple-choice item in a domain you know and understand. You can use the structured template to specify the topic, cognitive taxonomy, jurisdiction, response format, distractor rationale, and an optional scenario, or you can write your own custom instructions.

Judging process

Once your item is submitted, five large language models generate their own versions of the question. You are shown pairs of outputs and asked to judge which one is better, without knowing which model produced each version. This blind, pairwise comparison helps reduce bias and focuses the evaluation on the quality of the item itself.

Leaderboard

The results of these comparisons feed into the live Elo-based leaderboard. Over time, this produces a ranking of models based on the practical quality of the MCQ items they produce, as evaluated by the people who understand the content best.

Research lead

Tim Burnett

Founder, Test Community Network · Research Lead

tim@educationtech.consulting