Abstract
OBJECTIVES: In public health, access to research literature is critical to informing decision-making and to identify knowledge gaps. However, identifying relevant research is not a straightforward task since public health interventions are often complex, can have positive and negative impacts on health inequalities and are applied in diverse and rapidly evolving settings. We developed a "living" database of public health research literature to facilitate access to this information using Natural Language Processing tools. MATERIALS AND METHODS: Classifiers were identified to identify the study design (eg, cohort study or clinical trial) and relationship to factors that may be relevant to inequalities using the PROGRESS-Plus classification scheme. Training data were obtained from existing MEDLINE labels and from a set of systematic reviews in which studies were annotated with PROGRESS-Plus categories. RESULTS: Evaluation of the classifiers showed that the study type classifier achieved average precision and recall of 0.803 and 0.930, respectively. The PROGRESS-Plus classification proved more challenging with average precision and recall of 0.608 and 0.534. The FAIR database uses information provided by these classifiers to facilitate access to inequality-related public health literature. DISCUSSION: Previous work on automation of evidence synthesis has focused on clinical areas rather than public health, despite the need being arguably greater. CONCLUSION: The development of the FAIR database demonstrates that it is possible to create a publicly accessible and regularly updated database of public health research literature focused on inequalities. The database is freely available from https://eppi.ioe.ac.uk/eppi-vis/Fair. NETSCC ID NUMBER: NIHR133603.