An example counts as correct only when every target token is correct. A nearly right residue earns no partial credit for that example.
Hard ranks by the largest consecutively certified T on modulus identities seen during training, then by the corresponding Max T on unseen interpolation and extrapolation modulus sizes. Ties use accuracy at each profile's first uncertified rung. Every example in every rung through Max T must be exactly correct.
Hard is a hidden task evaluation and may change aspects of the recurrence itself; do not assume it is repeated squaring.