Abstract
Masked prediction learns by inferring missing variables from visible context. When does optimizing this conditional task recover the true joint data distribution? We study this question using an $\varepsilon$-identifiability modulus, which measures the worst-case joint-distribution error permitted by excess risk at most $\varepsilon$. For distributions with separated global modes, schedules retaining large visible contexts can permit substantial mode-weight errors at exponentially small excess risk. An exact information decomposition explains why: for a fixed mask, the loss penalizes only the mode-weight mismatch that remains unresolved by the visible context. For small mode-weight perturbations, the objective's sensitivity is proportional to residual mode uncertainty averaged over masks. Under joint masked-block log loss, low-visibility masks that retain mode uncertainty restore this sensitivity, while positive full-mask probability bounds joint-distribution error in terms of excess risk. We empirically validate these predictions through exact calculations and controlled stochastic optimization.