M. V. D. Prasad
Koneru Lakshmaiah Education Foundation

Published : 2 Documents Claim Missing Document
Claim Missing Document
Check
Articles

Found 2 Documents
Search

Computationally efficient handwritten Telugu text recognition Buddaraju Revathi; M. V. D. Prasad; Naveen Kishore Gattim
Indonesian Journal of Electrical Engineering and Computer Science Vol 34, No 3: June 2024
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/ijeecs.v34.i3.pp1618-1626

Abstract

Optical character recognition (OCR) for regional languages is difficult due to their complex orthographic structure, lack of dataset resources, a greater number of characters and similarity in structure between characters. Telugu is popular language in states of Andhra and Telangana. Telugu exhibits distinct separation between characters within a word, making a character-level dataset sufficient. With a smaller dataset, we can effectively recognize more words. However, challenges arise during the training of compound characters, which are combinations of vowels and consonants. These are considered as two or more characters based on associated vattus and dheerghams with the base character. To address this challenge, each compound character is encoded into a numerical value and used as input during training, with subsequent retrieval during recognition. The segmentation issue arises from overlapping characters caused by varying handwritten styles. For handling segmentation issues at the character level arising from handwritten styles, we have proposed an algorithm based on the language's features. To enhance word-level accuracy a dictionary-based model was devised. A neural network utilizing the inception module is employed for feature extraction at various scales, achieving word-level accuracy rates of 78% with fewer trainable parameters.
Computationally efficient ResNet based Telugu handwritten text detection Buddaraju Revathi; M. V. D. Prasad; Naveen Kishore Gattim
Bulletin of Electrical Engineering and Informatics Vol 13, No 6: December 2024
Publisher : Institute of Advanced Engineering and Science

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.11591/eei.v13i6.8170

Abstract

Optical character recognition (OCR) is a technological process that converts diverse document formats into editable and searchable data. Recognition of Telugu characters through OCR poses a challenge because of compound characters. Identifying handwritten Telugu text proves difficult due to the substantial number of characters, their similarities, and overlapping forms. To handle overlapping characters, we implemented a segmentation algorithm that efficiently separates these characters, consequently enhancing the model’s accuracy. Feature extraction is a crucial phase in recognizing a broader range of characters, especially those that are similar in appearance. So, we have employed a light weighted ResNet 34 model that effectively addresses these challenges and handles deep networks without declining accuracy as the network’s depth increases. We have achieved a word level recognition rate of 81.5%. In addition, the parameters required by the model are less when compared to its counterpart inception V1, making it computationally efficient.