Representativeness of data for model development
Institutions should analyse the representativeness of data in the case of statistical models and other mechanical methods used to assign exposures to grades or pools, as well as in the case of statistical default prediction models generating default probability estimates for individual obligors or facilities. Institutions should select an appropriate dataset for the purpose of model development to ensure that the performance of the model on the application portfolio, in particular its discriminatory power, is not significantly hindered by insufficient representativeness of data.
For the purposes of ensuring that the data used in developing the model for assigning obligors or exposures to grades or pools is representative of the application portfolio covered by the relevant model, as required in Article 174(c) of Regulation (EU) No 575/2013 and Article 40(2) of the RTS on IRB assessment methodology institutions should analyse the representativeness of the data at the stage of model development in terms of all of the following:
For the purpose of paragraph 21(a) institutions should analyse the segmentation of exposures and consider whether there were any changes to the scope of application of the considered model over the period covered by the data used in developing the model for assigning obligors or exposures to grades or pools. Where such changes were observed institutions should analyse the risk drivers relevant for the change of the scope of application of the model by comparing their distribution in the RDS before and after the change as well as with the distribution of those risk drivers in the application portfolio. For this purpose institutions should apply statistical methodologies such as cluster analysis or similar techniques to demonstrate representativeness. In the case of pooled models the analysis should be performed with regard to the part of the scope of the model that is used by an institution.
For the purpose of paragraph 21(b) institutions should ensure that the definition of default underlying the data used for model development is consistent over time and, in particular, that it is consistent with all of the following:
that adjustments have been made to achieve consistency with the current default definition where the default definition has been changed during the observation period;
that adequate measures have been adopted by the institution, where the model covers exposures in several jurisdictions having or having had different default definitions;
that the definition of default in each data source has been analysed separately;
that the definition of default used for the purposes of model development does not have a negative impact on the structure and performance of the rating model, in terms of risk differentiation and predictive power, where this definition is different from the definition of default used by the institution in accordance with Article 178 of Regulation (EU) No 575/2013.
For the purpose of paragraph 21(c) institutions should analyse the distribution and range of values of key risk characteristics of the data used in developing the model for risk differentiation in comparison with the application portfolio. With regard to LGD models, institutions should perform such analysis separately for non-defaulted and defaulted exposures.
Institutions should analyse the representativeness of the data in terms of the structure of the portfolio by relevant risk characteristics based on statistical tests specified in their policies to ensure that the range of values observed on these risk characteristics in the application portfolio is adequately reflected in the development sample. Where the application of statistical tests is not possible, institutions should carry out at least a qualitative analysis on the basis of the descriptive statistics of the structure of the portfolio, taking into account the possible seasoning effects referred to in Article 180(2)(f) of Regulation (EU) No 575/2013. When considering the results of this analysis, institutions should take into account the sensitivity of the risk characteristics to economic conditions. Material differences in the key risk characteristics between the data sample and the application portfolio should be addressed, for example by using another data sample or a subset of observations or by adequately reflecting these risk characteristics as risk drivers in the model.
For the purpose of paragraph 21(d) institutions should analyse whether, over the relevant historical observation period, there were significant changes in their lending standards or recovery policies or in the relevant legal environment, including changes in insolvency law, legal foreclosure procedures and any legal regulations related to realisation of collaterals, which may influence the level of risk or the distribution or ranges of the risk characteristics in the portfolio covered by the considered model. Where institutions observe such changes they should compare the data included in the RDS before and after the change of the policy. Institutions should ensure comparability of the current underwriting or recovery standards with those applied to the observations included in the RDS and used for model development.
Within the PD model the representativeness of data used in developing the model for risk differentiation does not require that the proportion of defaulted and non-defaulted exposures in this dataset be equal to the proportion of defaulted and non-defaulted exposures in the institution’s application portfolio. However, institutions should have a sufficient number of defaulted and non-defaulted observations in the development dataset and they should document the difference.