Ethics reference
What are the limits of a pre-hire assessment?
A pre-hire assessment is limited by the evidence supporting its intended use, the consistency of its scores, its performance across groups, and the information candidates receive before taking it. Validity, fairness, construct stability, and transparency must be reviewed together. A standardised process can reduce variation in administration without proving that every inference is correct or appropriate.
Validity is about use and evidence
Validity is not a permanent property attached to a test name. It concerns the interpretation and use of scores for a defined purpose. Evidence should connect the measured construct to a relevant criterion, such as a clearly specified aspect of job performance, in a population and role family that resemble the intended deployment. Evidence from another role or sample may inform a decision, but it does not automatically transfer.
Van Iddekinge, Lievens, and Sackett's review, Personnel selection: A review of ways to maximize validity, diversity, and the applicant experience covers ways to maximise selection validity, diversity, and applicant experience in Personnel Psychology (2023). Their framing supports a practical rule: a client should ask what the assessment predicts, for whom, and under which conditions before treating a score as decision-ready.
Fairness requires more than equal inputs
Giving every candidate the same stimulus is an administration control. It is not a complete fairness finding. Fairness review can examine score distributions, selection rates, error patterns, accessibility, and the downstream workflow in which a score is used. Different fairness definitions can conflict, so the selected definition and decision threshold should be documented rather than implied.
Raghavan, Barocas, Kleinberg, and Levy's study, Mitigating bias in algorithmic hiring: Evaluating claims and practices (FAccT, 2020) shows why disclosure of development and validation procedures is part of responsible evaluation. Mehrabi and colleagues' survey, A survey on bias and fairness in machine learning in ACM Computing Surveys (2021) further documents that bias can enter through data, labels, measurement choices, and deployment. Obscura's ethics and fairness record describes the product's stated controls and their limits.
Construct stability must be tested
A construct is stable when scores retain a defensible meaning across relevant occasions, groups, devices, and administration conditions. Stability does not require identical scores forever. It does require evidence that observed change is not primarily caused by avoidable changes in the instrument or scoring process. A construct that shifts with role framing, language, accessibility needs, or repeated exposure may need narrower interpretation.
Response timing and context can affect what a digital assessment records. Kyllonen and Zu's review, Use of response time for measuring cognitive ability in Journal of Intelligence (2016) illustrates why latency must be interpreted with attention to speed-accuracy trade-offs and model assumptions. A timestamp alone is not a psychological conclusion.
Candidate transparency sets a boundary
Candidates should know that an assessment is being used, what kind of information it records, the purpose of processing, and which organisation is responsible for the hiring decision. Access to raw responses may be different from access to proprietary scoring outputs, but the distinction should be stated plainly. Candidates also need a route for questions, corrections, accessibility support, and lawful data requests.
Transparency cannot turn weak evidence into strong evidence. It does make the decision process more inspectable and gives candidates information needed to exercise their rights. Obscura recommends that clients review the instrument's evidence and workflow before deployment, then monitor outcomes after deployment.
Last updated: 28 May 2026. Related record: how the analysis stages are separated.