Non-life insurance pricing is forward-looking, yet its principal cost signal, reported claims, is delayed by the occurrence-to-reporting process. The rating cell is treated as a stylised, homogeneous unit without renewal, lapse, expiry, or cohort dynamics; a fully annual-contract formulation would require these dynamics to be modelled jointly with the risk and claims processes. This paper develops a partially observed risk-sensitive control framework in which premium-sensitive exposure generates claims in a latent risk regime, while unreported claims form an atomic population governed by an age-structured transport equation. The numerical instance solved and validated in this research restricts the general model to a memoryless (one-phase) reporting process for tractability. It is best matched to lines with predominantly short-to-medium reporting tails, rather than to the most extreme long-tailed liability or cyber exposures the general model is designed to eventually accommodate. A finite-state reporting reservoir and nominal Bayesian filter provide the decision state, and compound Poisson–Gamma loss enters an entropic Bellman recursion through a closed-form exponential-tilting identity, checked against an independently coded Bellman-residual test. Every dynamic policy is benchmarked against alternatives matched at the same risk-sensitivity parameter and evaluated using common random numbers. This is a general feature of the results, not a single statistic:
is a standardised yardstick applied uniformly across strategies, while
at the policy’s own
is what that policy actually optimises, and the two need not agree. In 3000 out-of-model paths at
= 1.2, the delay-aware dynamic policy increases mean profit by EUR 0.148 million and a standardised
= 0.8 certainty equivalent by EUR 0.034 million relative to a matched static price. Its fifth percentile and TVaR 5%, by contrast, are lower by EUR 0.056 and 0.078 million. At the policy’s own optimisation level, however, the
= 1.2 certainty-equivalent difference is EUR −0.006 million, with a 95% interval reaching zero. The dynamic policy, therefore, does not clearly outperform the matched static price on the exact objective it was optimised to maximise. Matched comparisons attribute EUR 0.039–0.043 million of mean profit to reporting-delay modelling and EUR 0.017 million to dynamic continuation. Separating state observation from transition-law knowledge attributes EUR 0.161–0.182 million to observing the regime exactly, and a small, sign-changing EUR −0.010 to +0.003 million to knowing the true transition law itself. Across an eight-scenario misspecification stress suite, the dynamic policy’s mean-profit advantage over the matched static price is directionally robust in seven of eight scenarios, but it reverses sign under a +30% true-severity shock, indicating that this advantage is sensitive to substantial severity misspecification specifically. Grid, Bellman-residual, and out-of-model filter diagnostics indicate that state reconstruction is economically primary, while entropy risk sensitivity is a secondary overlay whose apparent benefit depends materially on which certainty-equivalent level and evaluation model are used to judge it.