Predictive validity, reliability, ecological validity: three key notions to judge a soft skills test. The vocabulary, the traps and the questions to ask before choosing an assessment tool.

Every test vendor promises a "reliable, validated" assessment. Few recruiters have the vocabulary to check what those words cover. That is a problem, because behind "scientifically validated" can hide very different realities: a questionnaire calibrated on thousands of assessments, or an in-house grid never confronted with the field. How do you judge the scientific quality of a soft skills test when you are not a psychometrician?
Predictive validity answers the only question that truly matters to a recruiter: are the test results linked to future behaviour on the job? A test can be pleasant to take, produce beautiful reports and predict nothing at all. Conversely, an austere measure can carry real information about future success.
Establishing predictive validity takes time and data: you must follow cohorts and compare initial scores with subsequently observed outcomes. It is research work, not a marketing argument. Hence a simple rule: when a vendor claims predictive validity, ask on which data, on which populations and over what time horizon it was established.
Reliability refers to the consistency of the measure: does the same person, under comparable conditions, obtain comparable results? A test whose results fluctuate strongly from one sitting to the next does not measure a competency, it measures noise.
This is the criterion on which many popular tools fail. Typologies that sort people into clear-cut categories are particularly exposed: close to one person in two changes MBTI profile when retaking the test a few weeks later. Our detailed analysis explains why the MBTI is not an HR decision tool.
A test can be reliable and predictive in the laboratory, yet far removed from real work situations. Ecological validity questions the resemblance between the assessed task and the targeted professional situation. Solving an abstract syllogism says little about the ability to settle a trade-off under pressure with incomplete data.
This is the central argument in favour of contextualised professional scenarios: the more the assessment situation resembles the work situation, the shorter the inference from test to job, and the more defensible the measure.
The volume trap. "Used by millions of people" is not a validity argument. Popularity measures diffusion, not the quality of the measure. Some of the most widely distributed tests in the world are also the most contested by research, as our comparison of the major personality tests shows.
The seniority trap. "70 years of existence" is not validation. Cattell's 16PF is a historically important, seriously constructed instrument; other tools from the same era survived through commercial habit more than scientific solidity. Age says a tool has lasted, not that it measures accurately.
The single-number trap. Beware of spectacular reliability percentages quoted without source or protocol. A psychometric claim is judged on consultable publications, described samples and acknowledged limitations, not on a round number in a brochure.
Rising Up's position rests on three verifiable choices.
A published framework. The 18 behavioural competencies of the Rising Up framework are publicly defined: operational definitions, families, observable behaviours. You cannot judge the validity of a measure if you do not know what it claims to measure; publishing the framework is the prerequisite of any serious audit. The assessment mechanics remain proprietary, the framework is available.
Situations rather than declarations. The Soft Skills Scan never asks about the competency itself. It places the person in concrete professional situations and analyses their responses, combining several assessment formats. This choice directly targets ecological validity: asking someone whether they are rigorous measures their self-image; observing their trade-offs in situation approaches the competency as exercised.
A continuous research programme. Recognised as Deeptech by the European Innovation Council (EIC) in 2023 and holder of the French CIR research accreditation, Rising Up conducts its work under the direction of its scientific team, in collaboration with researchers from the CNRS, ENS-PSL and Université Paris Cité. Validating an instrument is never a final state: it is a research programme that continues, is documented, and is refined.
Before choosing an assessment tool, demand written answers to these five questions:
A rigorous vendor answers precisely and owns the limitations. A fragile vendor answers "it is scientifically validated" and changes the subject.
These questions of method are not an expert's luxury. Behavioural skills predict professional success, contribute to it causally, and can be developed (research by James Heckman, Nobel laureate in economics). And the market has integrated these stakes: according to APEC, 41% of companies recruiting executives now formally assess soft skills, up 9 points in one year.
The more behavioural assessment weighs in hiring and development decisions, the more the quality of measurement becomes a matter of professional fairness: a candidate rejected on the basis of an invalid test suffers real harm. Choosing a tool whose scientific framework can be audited is not just good HR practice, it is a responsibility.
To see how these principles translate into a hiring process, discover the Rising Up recruitment test, designed as a decision aid backed by a public framework.
From recruitment At development skills, Rising Up provides reliable behavioral data to guide your HR and educational choices.