Change search
Link to record
Permanent link

Direct link
Ehrentraut, Claudia
Publications (4 of 4) Show all publications
Ehrentraut, C., Ekholm, M., Tanushi, H., Tiedemann, J. & Dalianis, H. (2018). Detecting hospital-acquired infections: A document classification approach using support vector machines and gradient tree boosting. Health Informatics Journal, 24(1), 24-42
Open this publication in new window or tab >>Detecting hospital-acquired infections: A document classification approach using support vector machines and gradient tree boosting
Show others...
2018 (English)In: Health Informatics Journal, ISSN 1460-4582, E-ISSN 1741-2811, Vol. 24, no 1, p. 24-42Article in journal (Refereed) Published
Abstract [en]

Hospital-acquired infections pose a significant risk to patient health, while their surveillance is an additional workload for hospital staff. Our overall aim is to build a surveillance system that reliably detects all patient records that potentially include hospital-acquired infections. This is to reduce the burden of having the hospital staff manually check patient records. This study focuses on the application of text classification using support vector machines and gradient tree boosting to the problem. Support vector machines and gradient tree boosting have never been applied to the problem of detecting hospital-acquired infections in Swedish patient records, and according to our experiments, they lead to encouraging results. The best result is yielded by gradient tree boosting, at 93.7percent recall, 79.7percent precision and 85.7percent F1 score when using stemming. We can show that simple preprocessing techniques and parameter tuning can lead to high recall (which we aim for in screening patient records) with appropriate precision for this task.

Keywords
clinical decision-making, databases and data mining, ehealth, electronic health records, secondary care
National Category
Information Systems
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-136583 (URN)10.1177/1460458216656471 (DOI)000424053900003 ()
Available from: 2016-12-12 Created: 2016-12-12 Last updated: 2022-03-23Bibliographically approved
Ehrentraut, C., Sundström, K. & Dalianis, H. (2015). Exploration of known and unknown early symptoms of cervical cancer and development of a symptom spectrum: Outline of a data and text mining based approach. In: John Krogstie, Gustaf Juel-Skielse, Vandana Kabilan (Ed.), Proceedings of the CAiSE-2015 Industry Track: co-located with 27th Conference on Advanced Information Systems Engineering (CAiSE 2015). Paper presented at 27th International Conference on Advanced Information Systems Engineering: CAiSE-2015 Industry Track, Stockholm, Sweden, 11 June, 2015 (pp. 34-44).
Open this publication in new window or tab >>Exploration of known and unknown early symptoms of cervical cancer and development of a symptom spectrum: Outline of a data and text mining based approach
2015 (English)In: Proceedings of the CAiSE-2015 Industry Track: co-located with 27th Conference on Advanced Information Systems Engineering (CAiSE 2015) / [ed] John Krogstie, Gustaf Juel-Skielse, Vandana Kabilan, 2015, p. 34-44Conference paper, Published paper (Refereed)
Abstract [en]

This position paper lays up the structure of some experiments to detect early symptoms of cervical cancer. We are using a large corpora of electronic patient records texts in Swedish from Karolinska University Hosptital from the years 2009-2010, where we extracted in total 1,660 patients with the diagnosis code C53. We used a Named Entity Recogniser called Clinical Entity Finder to detect the diagnosis and symptoms expressed in these clinical texts containing in total 2,988,118 words. We found 28,218 symptoms and diagnoses on these 1,660 patients. We present some initial findings, and discuss them and propose a set of experiments to find possible early symptoms or at least a spectrum or finger prints for early symptoms of cervical cancer.

Series
CEUR Workshop Proceedings, ISSN 1613-0073 ; 1381
National Category
Information Systems
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-119908 (URN)
Conference
27th International Conference on Advanced Information Systems Engineering: CAiSE-2015 Industry Track, Stockholm, Sweden, 11 June, 2015
Available from: 2015-11-11 Created: 2015-08-28 Last updated: 2022-02-23Bibliographically approved
Ehrentraut, C., Kvist, M., Sparrelid, E. & Dalianis, H. (2014). Detecting Healthcare-Associated Infections in Electronic Health Records: Evaluation of Machine Learning and Preprocessing Techniques. In: Proceedings of the 6th International Symposium on Semantic Mining in Biomedicine (SMBM 2014): . Paper presented at Sixth International Symposium on Semantic Mining in Biomedicine (SMBM 2014), Aveiro, Portugal, October 6-7, 2014 (pp. 3-10). University of Aveiro
Open this publication in new window or tab >>Detecting Healthcare-Associated Infections in Electronic Health Records: Evaluation of Machine Learning and Preprocessing Techniques
2014 (English)In: Proceedings of the 6th International Symposium on Semantic Mining in Biomedicine (SMBM 2014), University of Aveiro , 2014, p. 3-10Conference paper, Published paper (Refereed)
Abstract [en]

Healthcare-associated infections (HAI) are in- fections that patients acquire in the course of medical treatment. Being a severe pub- lic health problem, detecting and monitoring HAI in healthcare documentation is an impor- tant topic to address. Research on automated systems has increased over the past years, but performance is yet to be enhanced. The dataset in this study consists of 214 records obtained from a Point-Prevalence Survey. The records are manually classified into HAI and NoHAI records. Nine different preprocess- ing steps are carried out on the data. Two learning algorithms, Random Forest (RF) and Support Vector Machines (SVM), are applied to the data. The aim is to determine which of the two algorithms is more applicable to the task and if preprocessing methods will affect the performance. RF obtains the best performance results, yielding an F1 -score of 85% and AUC of 0.85 when lemmatisation is used as a preprocessing technique. Irrespec- tive of which preprocessing method is used, RF yields higher recall values than SVM, with a statistically significant difference for all but one preprocessing method. Regarding each classifier separately, the choice of preprocess- ing method led to no statistically significant improvement in performance results.

Place, publisher, year, edition, pages
University of Aveiro, 2014
National Category
Information Systems
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-108679 (URN)10.5167/uzh-98982 (DOI)
Conference
Sixth International Symposium on Semantic Mining in Biomedicine (SMBM 2014), Aveiro, Portugal, October 6-7, 2014
Available from: 2014-11-03 Created: 2014-11-03 Last updated: 2022-02-23Bibliographically approved
Ehrentraut, C., Ibrahim, O. & Dalianis, H. (2014). Text Analysis to support structuring and modelling a public policy problem: Outline of an algorithm to extract inferences from textual data. In: Gustaf Juell-Skielse (Ed.), DSV writers hut 2014: proceedings. Paper presented at DSV writers hut 2014, Åkersberga, Sweden, August 21-22, 2014. Stockholm: Department of Computer and Systems Sciences, Stockholm University
Open this publication in new window or tab >>Text Analysis to support structuring and modelling a public policy problem: Outline of an algorithm to extract inferences from textual data
2014 (English)In: DSV writers hut 2014: proceedings / [ed] Gustaf Juell-Skielse, Stockholm: Department of Computer and Systems Sciences, Stockholm University , 2014Conference paper, Published paper (Other academic)
Abstract [en]

Policy making situations are real-world problems that exhibit complexity in that they are composed of many interrelated problems and issues. To be effective, policies must holistically address the complexity of the situation rather than propose solutions to single problems. Formulating and understanding the situation and its complex dynamics, therefore, is a key to finding holistic solutions. Analysis of text based information on the policy problem, using Natural Language Processing (NLP) and Text analysis techniques, can support modelling of public policy problem situations in a more objective way based on domain experts’ knowledge and scientific evidence. The objective behind this study is to support modelling of public policy problem situations, using text analysis of verbal descriptions of the problem. We propose a formal methodology for analysis of qualitative data from multiple information sources on a policy problem to construct a causal diagram of the problem. The analysis process aims at identifying key variables, linking them by cause-effect relationships and mapping that structure into a graphical representation that is adequate for designing action alternatives, i.e., policy options. This study describes the outline of an algorithm used to automate the initial step of a larger methodological approach, which is so far done manually. In this initial step, inferences about key variables and their interrelationships are extracted from textual data to support a better problem structuring. A small prototype for this step is also presented.

Place, publisher, year, edition, pages
Stockholm: Department of Computer and Systems Sciences, Stockholm University, 2014
Series
Report Series / Department of Computer & Systems Sciences, ISSN 1101-8526 ; 14-019
Keywords
Public policy, problem structuring, qualitative analysis, Natural Language Processing, algorithm, inference extraction
National Category
Information Systems
Research subject
Computer and Systems Sciences
Identifiers
urn:nbn:se:su:diva-111107 (URN)978-91-637-7457-7 (ISBN)
Conference
DSV writers hut 2014, Åkersberga, Sweden, August 21-22, 2014
Available from: 2014-12-22 Created: 2014-12-22 Last updated: 2022-02-23Bibliographically approved
Organisations

Search in DiVA

Show all publications