Securing the Digital Exam: Integrity at Scale in Computer-Based Assessment

High-stakes assessment only has value if its results are trusted as reliable by candidates, institutions, regulators, and ultimately by the patients and public that medical professionals will serve. As digital assessment scales up and distributes across settings, deliberate design at every stage is required in order to maintain that trust.

The risks are real (and recent events in India show just how serious it can be when things go badly wrong, in this case with a paper exam), but they are manageable. The institutions handling security best are not those that have eliminated every threat — that is not achievable — but those that have thought carefully about where the vulnerabilities lie, built appropriate controls around them, and established the governance to respond when something goes wrong.

The Threat Landscape

Security in computer-based assessment covers several distinct categories of risk. Of these, the most significant are: item exposure and memorisation, where candidates share or recall exam content after sitting; impersonation, where the person taking the assessment is not the registered candidate; collusion during a sitting, where candidates communicate with each other or with outside sources; and the compromising of results data, through unauthorised access after the assessment has concluded.

Each of these requires a different response. Treating “security” as a single problem leads to solutions that address some risks while leaving others unexamined.

Item Banking and Exposure Risk

For high-stakes medical assessment, item exposure is arguably the most persistent and technically complex security challenge. Unlike a data breach, which is an event, item exposure is a process; it accumulates gradually and is very difficult to reverse. Once a question has circulated among candidates, the value of that item as a measure of genuine knowledge is compromised, and no amount of after-the-fact intervention can put the genie back in the bottle.

Managing exposure risk requires deliberate strategy at the item bank level. Pool size matters: a bank with sufficient depth to allow genuine rotation across cohorts and sittings reduces the probability that any individual item is seen repeatedly.

Planned item retirement — removing questions from active use before they become overexposed, even when they are still performing well psychometrically — is a discipline that well-run assessment programmes build into their governance calendar rather than treating as an emergency response.

Psychometric methods offer an additional layer of detection. Pre-knowledge of items tends to leave identifiable traces in response data: unusually high facility on specific questions that should require deliberation, response times that are implausibly fast for items of genuine difficulty, or score distributions that diverge anomalously from comparable cohorts. Monitoring for these patterns doesn’t prevent exposure, but can enable earlier detection of a problem and inform decisions about item retirement and investigation.

The governance dimension is as important as the technical one. Clear ownership of item compromise decisions — who has the authority to retire an item, on what evidence, and within what timeframe — is essential. Assessment programmes lacking this clarity risk allowing compromised items to remain in circulation longer than they should, because no one is empowered to act.

Controlling the Assessment Environment

One of the most direct vulnerabilities in digital assessment is the device itself. A candidate sitting an exam on a standard computer with unrestricted access to browsers, applications, and communication tools has, in effect, open access to external resources throughout the sitting.

Dedicated assessment applications that lock the host device during the exam address this at the platform level. Maxinity’s exam app prevents access to any other application or resource on the computer while the assessment is in progress. In addition it operates locally, so loss of internet connectivity does not disrupt the sitting. The security perimeter shifts from the room to the device, which is meaningful in both supervised and distributed settings.

Remote Invigilation — Capabilities and Honest Limits

Remote assessment became a practical necessity during the pandemic and has remained a feature of the landscape since, for reasons of access, scale, and flexibility. Understanding what remote invigilation can and cannot achieve is important for any institution making decisions about when and how to use it.

What it can achieve is significant: real-time identity verification at the point of entry; behaviour monitoring during the sitting; flagging of anomalous activity for human review; and a recorded audit trail that supports post-hoc investigation. AI-assisted invigilation, where automated detection identifies potential irregularities for human invigilators to review, makes it possible to maintain meaningful oversight across large numbers of simultaneous sittings in a way that purely human invigilation at scale cannot.

Maxinity’s Active Invigilation Module (AIM) enables this kind of real-time remote oversight, allowing invigilators to monitor candidates and intervene where necessary during a distributed sitting.

Here it should be acknowledged that remote invigilation can never fully replicate the controlled environment of a supervised examination hall because any remote setting involves variables outside institutional control. However, the appropriate response to this is not to avoid remote assessment but to make risk-based decisions about where it is suitable: the stakes involved, the candidate population, the nature of the content, and the compensating controls available.

Data Integrity and Auditability

Results data is the final output of the assessment process, and its integrity matters as much as the integrity of the sitting itself. Access controls that limit who can view, modify, or export results data, with clear audit trails of every action taken, are a baseline requirement. Encryption of data in transit and at rest, and clear protocols for how long data is retained and under what conditions it can be accessed, should be established before deployment rather than after the fact.

The audit trail deserves particular attention. In the event of a challenge to results — whether from a candidate, a regulator, or an internal review — institutions need to be able to demonstrate a complete and unbroken chain of custody: that the results accurately reflect what candidates submitted, that no unauthorised modification occurred, and that the process was carried out as documented. Institutions that have invested in auditability find that it also builds confidence with external stakeholders, including regulatory bodies that require evidence of assessment rigour.

Incident response planning is the necessary complement to these controls. No system is breach-proof, and the quality of an institution’s response when something goes wrong matters as much as its prevention measures. Clear protocols — who is notified, what is investigated, how candidates are told, and what remediation looks like — should be documented and tested before they are needed.

Integrity by Design

The institutions managing assessment security most effectively share a common characteristic: they do not treat security as a separate workstream added to an existing process. They build it in from the start, at item development, platform selection, invigilation strategy, data governance, and candidate communication.

When integrity is designed in rather than bolted on, it is also easier to demonstrate. Regulators, professional bodies, and the candidates themselves need confidence that results mean what they say. In high-stakes medical assessment, that confidence is the foundation on which everything else rests.

This is the second in Maxinity’s series on data and digital assessment in medical education. Our previous post examined the data and analytics potential of computer-based assessment at scale.