Our Accessibility & Inclusion Working Group has agreed on a first set of accessibility criteria for the FLIP+ Library, designed to be usable by institutions with no specialist, no team and no budget.
Ask most assessment organisations about accessibility, and you will hear about accommodations.
Extra time. A separate room. A rest break.
All necessary. All arriving after the item has already been written.
As assessment moves into digital environments, that sequence stops working. Accessibility has to sit inside design, item development and the interpretation of data. And the consequences are not only ethical; they land in measurement quality: missing responses, underestimated performance, comparability, validity.
A first set of criteria
Our Accessibility and Inclusion Working Group has agreed on a first set of item-level criteria for the FLIP+ Library, each with a definition, guidance on what to check, and worked examples.
They cover what determines whether an item is reachable at all: simple language unless the vocabulary is itself what is being measured; a logical order; images used only where they support the construct, with alternative text; a limit on information loaded into one item; explicit instructions rather than implicit expectations; freedom from stereotype; a flag where a cultural reference needs adaptation; contrast, scrolling, and a clear separation between source material and question.
Why accessibility criteria are built this way
Every criterion had to pass two tests: be supported by research and be checkable by someone without specialist training. That second test comes from evidence.
An international survey by the group's coordinator, Élodie Vezon, found practice strikingly uneven, with many institutions having no accessibility specialist, no dedicated team, and no budget for one. What they share is the same need: frameworks and guidelines they can actually use.
A checklist only an expert can apply would have helped the institutions that need it most.
It is not a box to tick that makes an item accessible. Judgement and training still matter. This is a first step and a starting point for a conversation inside an institution, not a substitute for one.
Part of a bigger question - item quality
One question now leads every conversation FLIP+ has: what makes an item a quality item, and how do you prove it?
Our answer is taking shape as a short record attached to each item: how it was made, what evidence supports it, whether AI was involved, whether a human reviewed it. Not a single pass-or-fail stamp, but each dimension graded separately and honestly where evidence is missing.
Accessibility is one of those dimensions. That is the point of this work: accessibility as part of item quality, not a separate conversation held elsewhere.
IAEA2026 in September in Toronto - building with the community
Élodie Vezon runs a pre-conference workshop at the 51st IAEA 2026 Annual Conference in Toronto: Trustworthy Data Begins with Accessible Design. Three case stations, namely visual access, language barriers, and neurodevelopmental barriers, all align around what an assessment is meant to measure, what the digital format gets in the way of, and which adaptations protect the construct.
We will present the first survey results there. If you are in Toronto, come and find us.
Coordinated by Élodie Vezon (DEPP, French Ministry of National Education), with members from Ireland, Lithuania, Brazil and the United Kingdom. FLIP+ Working Groups are open to member institutions. Write to us at info@flip-plus.org.
