Data Compilation Methods
explains data validation and editing, describes imputation for missing observations, develops population estimation through grossing up, and covers outlier identification and treatment
Data Compilation Methods
Data compilation refers to the operations performed on collected data to derive new information according to a given set of rules (statistical procedures), with a view to producing statistical outputs. Data compilation methods cover: (a) data validation and editing; (b) imputation of missing data; and (c) estimation of population characteristics. These methods address problems with collected data such as incomplete coverage, non-response, out-of-range responses, multiple responses, inconsistencies or contradictions, and invalid responses — problems that may stem from deficiencies in questionnaire design, lack of proper interviewer training, respondent error, and/or data-processing error. It is advisable to periodically generate reports on the frequency of each type of problem, to identify the main sources of error and adjust future data collection processes accordingly [IRES, Ch. VII, para. 7.61, PDF p. 110, 2018; see also IRIS 2008, footnote 62, for more on data-compilation techniques].
Data validation and editing
Data validation and editing is essential to assuring the quality of collected data. It refers to the systematic examination of data collected from respondents to identify, and eventually modify, inadmissible, inconsistent and highly questionable or improbable values, according to predetermined rules. Validation criteria must clearly and systematically confirm whether data satisfy requirements of completeness, integrity, arithmetic consistency and congruence, and guarantee overall data quality. Criteria are established by the statistical authority according to the nature of the data and analysis of the variables of interest, taking into account magnitude, structure, trends, relationships, causalities, interdependencies and possible response ranges [IRES, Ch. VII, para. 7.62, PDF p. 110, 2018].
Any arbitrary alteration of data should not be allowed; changes to collected data must be based on the relationship between variables and response values. Appropriate response ranges for each question, and the congruence that must exist between responses to related questions, must be established to prevent out-of-range responses and inconsistencies. For example, checking that the sum of available supplies equals the sum of recorded uses is an important validation criterion — including for routine questionnaires directed to the energy industries [IRES, Ch. VII, para. 7.63, PDF p. 110, 2018].
Because validation and editing can be an expensive part of the survey process, attention should focus on the most important areas and issues. Many survey responses have minimal impact on final results, so correcting errors in such responses may not be worth the effort; the responses with the greatest impact on final results should be identified before the validation process begins, so that resources can be properly allocated [IRES, Ch. VII, para. 7.64, PDF p. 111, 2018].
Data imputation
Imputation refers to replacing one or more erroneous responses or non-responses with plausible, internally consistent values to produce a complete data set. It is used to estimate missing data values — for example, when a respondent has answered only part of the relevant questions, or when answers are not logically correct. Imputation methods range from simple and intuitive to complex statistical procedures [IRES, Ch. VII, para. 7.65, PDF p. 111, 2018].
The choice of imputation method depends on the objective of the analysis and the type of missing data; no method is superior to all others in every circumstance, and most imputation systems use a mix of methods [IRES, Ch. VII, para. 7.66, PDF p. 111, 2018; see also IRIS, Ch. VI.B.2, footnote 63, for imputation options for item vs. unit non-response]. Desirable properties of all imputation processes are:
(a) Imputed records should closely resemble the missing or failed-edit record, retaining as much respondent data as possible — the number of imputed data items should be kept to a minimum;
(b) Imputed records should satisfy all edit checks;
(c) Imputed values should be flagged, and the methods and sources used for imputation described in the metadata
[IRES, Ch. VII, para. 7.66(a)–(c), PDF p. 111, 2018].
Recommendation (7.67): it is recommended that compilers of energy statistics use imputation as necessary, with appropriate methods consistently applied, and that these methods comply with the general requirements set out in international recommendations for other domains of economic statistics, such as the International Recommendations for Industrial Statistics (UN 2009b) [IRES, Ch. VII, para. 7.67, PDF pp. 111–112, 2018] — recommendations tracker row VII/7.67.
Estimation of population characteristics
Grossing up
After data have been validated, edited and imputed for non-response and erroneous responses, special procedures — grossing up procedures — are applied to sample values to estimate the required characteristics of the total population. These procedures raise the sample value by a factor based on the sampling fraction, to obtain data levels for the full sample frame population. In some cases, depending on relationships with other variables for which data exist, more sophisticated statistical techniques may be used [IRES, Ch. VII, para. 7.68, PDF p. 112, 2018].
Recommendation (7.68): as the application of estimation procedures is a complex undertaking, it is recommended that specialist expertise always be sought for this task [IRES, Ch. VII, para. 7.68, PDF p. 112, 2018] — recommendations tracker row VII/7.68.
Treatment of outliers
The treatment of outliers is an important estimation consideration, particularly in energy statistics. Outliers are reported data that are correct but unusual, in the sense that they do not represent the sampled population and may therefore distort estimates. If the sampling weight is large and the unadjusted outlier value is included in the sample, the final estimate will be inappropriately large and unrepresentative, driven by one extreme value. The simplest treatment is to reduce the outlier’s weight in the sample so that it represents only itself; alternatively, statistical techniques can calculate a more appropriate weight for the outlier unit. Details of outlier treatment should be provided in the metadata [IRES, Ch. VII, para. 7.69, PDF p. 112, 2018].
Related
- Statistical Data Sources and Administrative Data Sources — the data these compilation methods are applied to
- International Recommendations for Industrial Statistics — IRIS 2008, cited as the reference standard for imputation methods and for further detail on compilation techniques generally
Source material
This page is a cited synthesis. Read the cleaned source used for it: