THE COMPARISONS OF OCR TOOLS: A CONVERSION CASE IN THE MALAYSIAN HANSARD CORPUS DEVELOPMENT

Authors

  • Anis Nadiah Che Abdul Rahman Faculty of Social Sciences and Humanities, Universiti Kebangsaan Malaysia, Selangor
  • Imran Ho Abdullah Faculty of Social Sciences and Humanities, Universiti Kebangsaan Malaysia, Selangor
  • Intan Safinaz Zainuddin Faculty of Social Sciences and Humanities, Universiti Kebangsaan Malaysia, Selangor
  • Azhar Jaludin Faculty of Social Sciences and Humanities, Universiti Kebangsaan Malaysia, Selangor

DOI:

https://doi.org/10.24191/mjoc.v4i2.5626

Keywords:

Optical Character Recognition, PDF to text converter, Malay text converter, Corpus development, Malaysian Hansard Corpus

Abstract

Optical Character Recognition (OCR) is a tool in computational technology that allows a recognition of printed characters by manipulating photoelectric devices and computer software. It runs by converting images or texts that are scanned beforehand into machine-readable and editable texts. There are a various numbers of OCR tools in the market for commercial and research use, which are obtainable for free or restrained with purchases. An OCR tool is able to enhance the accuracy of the results which as well relies on pre-processing and subdivision of algorithms. This study intends to investigate the performances of OCR tools in converting the Parliamentary Reports of Hansard Malaysia for developing the Malaysian Hansard Corpus (MHC). By comparing four OCR tools, the study has converted ten reports of Parliamentary Reports which contains a number of 62 pages to see the conversion accuracy and error rate of each conversion tool. In this study, all of the tools are manipulated to convert Adobe Portable Document Format (PDF) files into Plain Text File (txt). The objective of this study is to give an overview based on accuracy and error rate of how each OCR tools essentially works and how it can be utilized to provide assistance towards corpus building. The study indicates that each tool possesses a variety of accuracy and error rates to convert the whole documents from PDF into txt or plain text files. The study proposes that a step of corpus building can be made easier and manageable when a researcher understands the way an OCR tool works in order to choose the best OCR tool prior to the outset of the corpus development.

References

ABBYY. (2018). ABBY FineReader Version 14 for Windows. Retrieved from https://www.abbyy.com/en-apac/finereader/

Afli, H., Barrault, L., & Schwenk, H. (2016). OCR Error Correction Using Statistical Machine Translation. International Journal of Computational Linguistics and Applications. 7(1), pp. 175–191.

Alexandrov, V. (2003). Error Evaluation and Applicability of OCR Systems. International Conference on Computer Systems and Technologies. Retrieved April 13, 2018 from http://ecet.ecs.uni-ruse.bg/cst/docs/proceedings/S3/III-10.pdf

Davenport, H.T., & Kirby, J. (2016). Just How Smart are Smart Machines? MITSloan Management Review. 57(3). pp 20 – 26. Retrieved 27 March 2019 from https://pdfs.semanticscholar.org/a9d8/0b09f21d9d0306766d2c3ba2ce49b4b2b95b.pdf

Diachronic. (2018). In Cambridge Dictionary. Retrieved from https://dictionary.cambridge.org/dictionary/english/diachronic

Heliński, M, Kmieciak , M., & Parkoła , T. (2012). Report on the comparison of Tesseract and ABBYY FineReader OCR engines. IMPACT. Retrieved July 20, 2017, from http://lib.psnc.pl/dlibra/docmetadata?from=rss&id=358

Herceg,P., Huyck, B., Johnson, C., Van Guilder, L. and Kundu, A. (2005). Optimizing OCR accuracy for bi-tonal, noisy scans of degraded Arabic documents. Visual Information Processing XIV. 179-187. Retrieved April 4, 2019 from https://pdfs.semanticscholar.org/8ed7/76c183ba07bccd47de941077f1e3b18f962a.pdf

Imran, H.A., Anis Nadiah C.A.R., Azhar, J. (2018). Malaysian Hansard Corpus. Universiti Kebangsaan Malaysia.

Iris, S.A. (2018). ReadIris Version 17 for Windows

Islam, N., Islam, Z., & Noor, N. (2017). A Survey on Optical Character Recognition System. Journal of Information & Communication Technology. 10(2), pp.1-4

Kimari, K. (2018, January 03). 8 best OCR software for Windows 10. Retrieved April 12, 2018, from https://windowsreport.com/ocr-software-windows-10/

Media4x. (2019). PDF to Text. Retrieved from https://pdftotext.com

Nield, D., DeMuro, J. & Turner, B. (2019). Best OCR Software of 2019: scan and archive your documents to PDF. TechRadar. Retrieved 10 October 2019 from https://www.techradar.com/best/best-ocr-software

Somer, M.(2010). Media Values and Democratization: What Unites and What Divides ReligiousConservative and Pro-Secular Elites? Turkish Studies. 11(4), pp 555-577

Plaintext. (2019) In Merriam-Webster, Merriam-Webster. Retrieved from www.merriam-webster.com/dictionary/plaintext.

Richter C., Wickes, M., Beser, D. & Marcus, M. (2018). Low-resource Post Processing of Noisy OCR Output for Historical Corpus Digitisation. Proceedings of the Eleventh International Conference on Language Resources and Evaluation. (2331-2339). Miyazaki: European Languages Resources Association (ELRA)

Riffe, D., Lacy, S., & Fico, F. G. (1998). Analyzing Media Messages: Using Quantitative Content Analysis in Research. London: Lawrence Erlbaum Associates.

Riffe, D., Lacy, S., & Fico, F. (2005). Analyzing Media Messages: Using Quantitative Analysis in Research. Mahwah, NJ: Lawrence ErlbaumAssociates.

4Videosoft Studio. (2018). PDF Converter Ultimate for Windows. Retrieved

from https://www.4videosoft.com/pdf-converter-ultimate.html

Published

2019-12-01

How to Cite

Anis Nadiah Che Abdul Rahman, Imran Ho Abdullah, Intan Safinaz Zainuddin, & Azhar Jaludin. (2019). THE COMPARISONS OF OCR TOOLS: A CONVERSION CASE IN THE MALAYSIAN HANSARD CORPUS DEVELOPMENT. Malaysian Journal of Computing, 4(2). https://doi.org/10.24191/mjoc.v4i2.5626