Record linkage for routinely collected health data in an African health information exchange

被引:7
|
作者
Mutemaringa, Themba [1 ,2 ,3 ]
Heekes, Alexa [1 ,2 ]
Smith, Mariette [1 ,2 ]
Boulle, Andrew [1 ,2 ]
Tiffin, Nicki [4 ,5 ]
机构
[1] Cape Govt Hlth, Prov Hlth Data Ctr, Hlth Intelligence Directorate, Cape Town, Western Cape, South Africa
[2] Univ Cape Town, Ctr Infect Dis Epidemiol & Res, Sch Publ Hlth & Family Med, Cape Town, South Africa
[3] Univ Cape Town, Computat Biol Div Integrat Biomed Sci Dept, Cape Town, South Africa
[4] Wellcome Ctr Infect Dis Res Africa, Fac Hlth Sci, Cape Town, South Africa
[5] Western Cape, South African Natl Bioinformat Inst, Cape Town, South Africa
基金
美国国家卫生研究院; 英国惠康基金; 比尔及梅琳达.盖茨基金会;
关键词
health information exchange; data linkage; global South; routine health data; Africa; South Africa;
D O I
10.23889/ijpds.v8i1.1771
中图分类号
R19 [保健组织与事业(卫生事业管理)];
学科分类号
摘要
Introduction The Patient Master Index (PMI) plays an important role in management of patient information and epidemiological research, and the availability of unique patient identifiers improves the accuracy when linking patient records across disparate datasets. In our environment, however, a unique identifier is seldom present in all datasets containing patient information. Quasi identifiers are used to attempt to link patient records but sometimes present higher risk of over-linking. Data quality and completeness thus affect the ability to make correct linkages. Aim This paper describes the record linkage system that is currently implemented at the Provincial Health Data Centre (PHDC) in the Western Cape, South Africa, and assesses its output to date. Methods We apply a stepwise deterministic record linkage approach to link patient data that are routinely collected from health information systems in the Western Cape province of South Africa. Variables used in the linkage process include South African National Identity number (RSA ID), date of birth, year of birth, month of birth, day of birth, residential address and contact information. Descriptive analyses are used to estimate the level and extent of duplication in the provincial PMI. Results The percentage of duplicates in the provincial PMI lies between 10% and 20%. Duplicates mainly arise from spelling errors, and surname and first names carry most of the errors, with the first names and surname being different for the same individual in approximately 22% of duplicates. The RSA ID is the variable mostly affected by poor completeness with less than 30% of the records having an RSA ID. The current linkage algorithm requires refinement as it makes use of algorithms that have been developed and validated on anglicised names which might not work well for local names. Linkage is also affected by data quality-related issues that are associated with the routine nature of the data which often make it difficult to validate and enforce integrity at the point of data capture.
引用
收藏
页数:13
相关论文
共 50 条
  • [41] Cardiovascular disease incidence rates: a study using routinely collected health data
    Ramroth, Johanna
    Shakir, Rebecca
    Darby, Sarah C.
    Cutter, David J.
    Kuan, Valerie
    CARDIO-ONCOLOGY, 2023, 9 (01)
  • [42] Using routinely collected health record data for the earlier detection of heart failure with preserved ejection fraction: FIND-HFpEF
    Nadarajah, Ramesh
    Nakao, Yoko M.
    Wu, Jianhua
    Gale, Chris P.
    EUROPEAN HEART JOURNAL, 2023, 44 (33) : 3113 - 3115
  • [43] An Overview of Record Linkage Methods: Applications and Perspective on Health Data
    Bounebache, Said Karim
    Quantin, Catherine
    Benzenine, Eric
    Obozinski, Guillaume
    Rey, Gregoire
    JOURNAL OF THE SFDS, 2018, 159 (03): : 79 - 123
  • [44] Using Routinely Collected Administrative Data in Public Health Research: Geocoding Alcohol Outlet Data
    Richard J. Fry
    Sarah E. Rodgers
    Jennifer Morgan
    Scott Orford
    David L. Fone
    Applied Spatial Analysis and Policy, 2017, 10 : 301 - 315
  • [45] Using Routinely Collected Administrative Data in Public Health Research: Geocoding Alcohol Outlet Data
    Fry, Richard J.
    Rodgers, Sarah E.
    Morgan, Jennifer
    Orford, Scott
    Fone, David L.
    APPLIED SPATIAL ANALYSIS AND POLICY, 2017, 10 (02) : 301 - 315
  • [46] ANNUAL HEALTH INSURANCE TREATMENT COST OF ALLERGIC ASTHMA BASED ON ROUTINELY COLLECTED HEALTH CARE FINANCING DATA
    Ponusz, R.
    Endrei, D.
    Elmer, D.
    Nemeth, N.
    Horvath, L.
    Csakvari, T.
    Sebestyen, A.
    Boncz, I
    VALUE IN HEALTH, 2020, 23 : S351 - S351
  • [47] Health Information Exchange: A Novel Re-linkage Intervention in an Urban Health System
    Sharp, Joseph
    Angert, Christine D.
    McConnell, Tyania
    Wortley, Pascale
    Pennisi, Eugene
    Roland, Lisa
    Mehta, C. Christina
    Armstrong, Wendy S.
    Shah, Bijal
    Colasanti, Jonathan A.
    OPEN FORUM INFECTIOUS DISEASES, 2019, 6 (10):
  • [48] ANNUAL HEALTH INSURANCE TREATMENT COST OF OTHER CATARACT BASED ON ROUTINELY COLLECTED HEALTH CARE FINANCING DATA
    Ponusz, R.
    Endrei, D.
    Kovacs, D.
    Nemeth, N.
    Molics, B.
    Danku, N.
    Csakvari, T.
    Boncz, I
    VALUE IN HEALTH, 2022, 25 (01) : S104 - S104
  • [49] ANNUAL HEALTH INSURANCE TREATMENT COST OF SENILE CATARACT BASED ON ROUTINELY COLLECTED HEALTH CARE FINANCING DATA
    Ponusz, R.
    Endrei, D.
    Kovacs, D.
    Boncz, I
    VALUE IN HEALTH, 2022, 25 (01) : S103 - S103
  • [50] Data Encryptions Techniques for Electronic Health Record Exchange
    Shin, David
    Sahama, Tony
    Kim, Steve
    Kim, Ji-Hong
    INTERNATIONAL PERSPECTIVES IN HEALTH INFORMATICS, 2011, 164 : 392 - 396