PenFit is tested for structural bias before the results are used to interpret individual or population preferences.
Important: these are not two million real individuals. PenFit ran 2,000,000 symmetric synthetic questionnaire profiles through the candidate engine to test whether the model itself favours one retirement strategy when the underlying responses are deliberately neutral and balanced.
2,000,000synthetic response profiles stress-tested
6 / 6retirement strategies can rank first
57.9neutral score for all six — exact tie
11.6–21.2%top-frequency range in the symmetric test
82.3–90.3highest winning alignment scores across the six strategies
Symmetric stress-test results
| Retirement approach | Top frequency | Highest winning alignment | Check |
| Flexible Income | 17.4% | 82.3 | PASS |
| Secure Income | 21.2% | 90.3 | PASS |
| Target Income | 14.2% | 85.8 | PASS |
| Flexible and Secure Income Mix | 19.8% | 86.0 | PASS |
| Flexible Income then Secure Income Later | 11.6% | 84.8 | PASS |
| Flexible Income with Secure Income Build-Up | 15.8% | 86.0 | PASS |
The purpose is not to force each strategy to win 16.7% of the time or to give every strategy the same maximum. Different retirement designs can legitimately occupy different areas of preference space. The test is designed to identify structural dominance or a strategy that is effectively unable to win.
What PenFit changed to reduce structural bias
1. Balanced questionnaire structureEach of the 12 retirement values appears twice and every answer uses the same +2 / +1 / 0 / −1 / −2 evidence scale.
2. Common neutral starting pointThe strategy-specific starting advantage is removed. An all-A/B response now gives every strategy exactly the same 57.9 score.
3. Normalised score responsivenessStrategies naturally move by different amounts when answers change. PenFit normalises that dispersion so a strategy does not win simply because its score has a wider natural range.
4. No artificial equalisationPenFit does not add strategy bonuses, force equal win rates or turn the top result into 100%. Genuine differences and close trade-offs remain visible.
Double-counting and sensitivity check
Correlated-characteristic test passed. PenFit repeated the calibration after capping overlapping security and flexibility characteristic families. All six strategies still became the top result; top frequencies remained about 12.6%–21.8%, and maximum winning scores remained within about ±4.8 points of their mean.
13 leave-one-characteristic-out checks passed. Removing each retirement characteristic in turn did not make any strategy impossible to select. No single characteristic appears to be driving the result.
Next behavioural control: the production Individual questionnaire should randomise whether each statement appears as A or B, with the scoring direction reversed automatically. This reduces systematic first/second-position effects without changing what is being asked.
Audit conclusion: the candidate calibration removes the structural preference identified in the earlier prototype and remains robust to the correlated-characteristic sensitivity checks. That is evidence of model neutrality under the conditions tested — not proof that any model is completely free from bias. Real-person testing and ongoing monitoring remain part of the validation process.