cover
Contact Name
Aji Prasetya Wibawa
Contact Email
keds.journal@um.ac.id
Phone
+62818539333
Journal Mail Official
keds.journal@um.ac.id
Editorial Address
Semarang St. No. 5, Malang, East Java, 65145, Indonesia
Location
Kota malang,
Jawa timur
INDONESIA
Knowledge Engineering and Data Science
ISSN : -     EISSN : 25974637     DOI : http://dx.doi.org/2597-4637
The journal welcomes experimental and theoretical findings on data science and knowledge engineering along with their applications to real-life situations.
Articles 117 Documents
Similarity Identification of Large-scale Biomedical Documents using Cosine Similarity and Parallel Computing Wibowo, Merlinda; Quix, Christoph; Hussien, Nur Syahela; Yuliansyah, Herman; Adhinata, Faisal Dharma
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Document similarity computation is an important research topic in information retrieval, and it is a crucial issue for automatic document categorization. The similarity value is between 0 and 1, then the closest value to 1 is represented both documents is considered more relevant, vice versa. However, the large scale of textual information has created the problem of finding the relevance level between documents. Therefore, the relevance between mesh heading text in the PubMed documents is higher than the relevance of the abstract text in the PubMed documents. Furthermore, parallel computing is implemented to speed up the large-scale documents similarity identification process that automatically calculates in the PubMed application. The execution time of mesh heading is 15.447 seconds, and the timely execution of abstract is 74.191 seconds. The execution time of mesh heading is higher than abstract because abstract contains more words than mesh heading. This study has successfully identified the similarity between large-scale biomedical documents of the PubMed documents that implemented a cosine similarity algorithm. The result has shown that the cosine similarity of the mesh heading texts is higher than the abstract text in the form of a graph and table shown in the PubMed application. The cosine similarity is useful to measure the similarity between documents based on the TF*IDF calculation result.
Recognition of Handwritten Javanese Script using Backpropagation with Zoning Feature Extraction Handayani, Anik Nur; Herwanto, Heru Wahyu; Chandrika, Katya Lindi; Arai, Kohei
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Backpropagation is part of supervised learning, in which the training process requires a target. The resulting error is transmitted back to the units below in its training process. Backpropagation can solve complicated problems because it consumes less memory than other algorithms. In addition, it also can produce solutions with a low error rate while executing less time. In image pattern recognition, backpropagation can be utilized for cultural preservation in many places worldwide, including Indonesia. It is used to recognize picture patterns in Javanese script writings. This study concluded that feature extraction approaches, zoning, and backpropagation could be utilized to distinguish handwritten Javanese characters. The best accuracy is attained at 77.00%, with the network architecture comprising 64 input neurons, 40 hidden neurons, a learning rate of 0.003, a momentum of 0.03, and an iteration of 5000.
A Comparative Study of Machine Learning-based Approach for Network Traffic Classification Trang, Kien; Nguyen, An Hoang
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Internet usage has increased rapidly and become an essential part of human life, corresponding to the rapid development of network infrastructure in recent years. Thus, protecting users’ confidential information when joining the global network becomes one of the most significant considerations. Even though multiple encryption algorithms and techniques have been applied in different parties, including internet providers, and web hosting, this situation also allows the hacker to attack the network system anonymously. Therefore, the significance of classifying network data streams to improve network system quality and security is attracting increasing study interests. This work introduces a machine learning-based approach to find the most suitable training model for network traffic classification tasks. Data pre-processing is first applied to normalize each feature type in the dataset. Different machine learning techniques, including k-Nearest Neighbors (KNN), Artificial Neural Network (ANN), and Random Forest (RF), are applied based on the normalized features in the classification phase. An open-access dataset ISCXVPN2016 is applied for this research, which includes two types of encryption (VPN and Non-VPN) and seven classes of traffic categories classes. Experimental results on the open dataset have shown that the proposed models have reached a high classification rate – over 85% in some cases, in which the RF model obtains the most refined results among the three techniques.
CNN based Face Recognition System for Patients with Down and William Syndrome Setyati, Endang; Az, Suharyono; Hudiono, Subroto Prasetya; Kurniawan, Fachrul
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Down syndrome, also known as trisomy genetic condition, is a genetic disorder that affects many people. Williams syndrome is a hereditary disorder that can affect anyone at birth. It marks medical and cognitive issues, such as cardiovascular illness, developmental delays, and learning impairments. This is accompanied by exceptional verbal abilities, a gregarious attitude, and a passion for music. Down syndrome and William Syndrome are both genetic illnesses. However, it can be distinguished from the arrangement of chromosome 21. Down syndrome and William syndrome can also be identified by recognizing faces, or facial characteristics, such as observing particular facial features. Therefore, this research develops Convolutional Neural Network (CNN) architectures to recognize Down syndrome and William syndrome using a facial recognition approach. A total of 480 facial photos were used in the study, with 390 images used for training data and 90 images used for testing data. The identification class is divided into three categories, Down syndrome, William syndrome, and normal. There are 160 photos in each patient class. This research presents two CNN architectures using a grayscale image of 256×256 pixels. The first CNN architecture comprises 12 layers, while the second comprises 15 layers. The average accuracy results with 12 layers were 91% by attempting to train and test six times. With 15 layers, the average accuracy value is 89%. In comparison, the first architecture has the highest accuracy value
Non-Gaussian Analysis of Herbarium Specimen Damageto Optimize Specimen Collection Management Yaman, Aris; Kartika, Yulia Aris; Indrawati, Ariani; Akbar, Zaenal; Manik, Lindung P.; Wardani, Wita; Djarwaningsih, Tutie; Mahendra, Taufik; Saleh, Dadan R.
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Damage to specimen collections occurs in practically every herbarium across the world. Hence, some precautions must be taken, such as investigating the factors that cause specimen damage in their collections and evaluating their herbarium collection handling and usage policy. However, manual investigation of the causes of herbarium collection damage requires a lot of effort and time. Only a few studies have attempted to investigate the causes of herbarium collection damage. So far, the non-gaussian approach to detecting the causes of damage to herbarium specimens has not been studied before. This study attempted to explore the effect of species type, time, location, storage, and remounting status on the level of damage to herbarium specimens, especially those in the genus Excoecaria. Gaussian modeling is not good enough to model the counted data phenomenon (the amount of damage to herbarium specimens). Negative binomial regression (NBR) provides a better model when compared to generalized Poisson regression and ordinary Gaussian regression approaches. NBR detects non-uniformity in the storage process, causing damage to herbarium specimens. Natural damage to herbarium specimens is caused by differences in species and the origin of specimens.
Parallel Approach of Adaptive Image Thresholding Algorithm on GPU Prahara, Adhi; Pranolo, Andri; Anwar, Nuril; Mao, Yingchi
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Image thresholding is used to segment an image into background and foreground using a given threshold. The threshold can be generated using a specific algorithm instead of a pre-defined value obtained from observation or experiment. However, the algorithm involves per pixel operation, histogram calculation, and iterative procedure to search the optimum threshold that is costly for high-resolution images. In this research, parallel implementations on GPU for three adaptive image thresholding methods, namely Otsu, ISODATA, and minimum cross-entropy, were proposed to optimize their computational times to deal with high-resolution images. The approach involves parallel reduction and parallel prefix sum (scan) techniques to optimize the calculation. The proposed approach was tested on various sizes of grayscale images. The result shows that the parallel implementation of three adaptive image thresholding methods on GPU achieves 4-6 speeds up compared to the CPU implementation, reducing the computational time significantly and effectively dealing with high resolution images.
Social Distancing Monitoring System using Deep Learning Ismail, Amelia Ritahani; Affendy, Nur Shairah Muhd; Puzi, Asmarani Ahmad
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

COVID-19 has been declared a pandemic in the world by 2020. One way to prevent COVID-19 disease, as the World Health Organization (WHO) suggests, is to keep a distance from other people. It is advised to stay at least 1 meter away from others, even if they do not appear to be sick. The reason is that people can also be the virus carrier without having any symptoms. Thus, many countries have enforced the rules of social distancing in their Standard Operating Procedure (SOP) to prevent the virus spread. Monitoring the social distance is challenging as this requires authorities to carefully observe the social distancing of every single person in a surrounding, especially in crowded places. Real-time object detection can be proposed to improve the efficiency in monitoring the social distance SOP inspection. Therefore, in this paper, object detection using a deep neural network is proposed to help the authorities monitor social distancing even in crowded places. The proposed system uses the You Only Look Once (YOLO) v4 object detection models for the detection. The proposed system is tested on the MS COCO image dataset with a total of 330,000 images. The performance of mean average precision (mAP) accuracy and frame per second (FPS) of the proposed object detection is compared with Faster Region-based Convolutional Neural Network (R-CNN) and Multibox Single Shot Detector (SSD) model. Finally, the result is analyzed among all the models.
Automatic 3D Cranial Landmark Positioning based onSurface Curvature Feature using Machine Learning Suputra, Putu Hendra; Sensusiati, Anggraini Dwi; Artaria, Myrtati Dyah; Verkerke, Gijsbertus Jacob; Yuniarno, Eko Mulyanto; Purnama, I Ketut Eddy
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

Cranial anthropometric reference points (landmarks) play an important role in craniofacial reconstruction and identification. Knowledge to detect the position of landmarks is critical. This work aims to locate landmarks automatically. Landmarks positioning using Surface Curvature Feature (SCF) is inspired by conventional methods of finding landmarks based on morphometrical features. Each cranial landmark has a unique shape. With the appropriate 3D descriptors, the computer can draw associations between shapes and landmarks using machine learning. The challenge in classification and detection in three-dimensional space is to determine the model and data representation. Using three-dimensional raw data in machine learning is a serious volumetric issue. This work uses the Surface Curvature Feature as a three-dimensional descriptor. It extracts the local surface curvature shape into a projection sequential value (depth). A machine learning method is developed to determine the position of landmarks based on local surface shape characteristics. Classification is carried out from the top-n prediction probabilities for each landmark class, from a set of predictions, then filtered to get pinpoint accuracy. The landmark prediction points are hypothetically clustered in a particular area, so a cluster-based filter is appropriate to isolate them. The learning model successfully detected the landmarks, with the average distance between the prediction points and the ground truth being 0.0326 normalized units. The cluster-based filter is implemented to increase accuracy compared to the ground truth. Thus, SCF is suitable as a 3D descriptor of cranial landmarks.
A Comparison of Machine Learning Models to Prioritise Emailsusing Emotion Analysis for Customer Service Excellence Chuttur, Mohammad Yasser; Parianen, Yashinee
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

There has been little research on machine learning for email prioritization for customer service excellence. To fill this gap, we propose and assess the efficacy of various machine learning techniques for classifying emails into three degrees of priority: high, low, and neutral, based on the emotions inherent in the email content. It is predicted that after emails are classified into those three categories, recipients will be able to respond to emails more efficiently and provide better customer service. We use the NRC Emotion Lexicon to construct a labeled email dataset of 517,401 messages for our proposal. Following that, we train and test four prominent machine learning models, MNB, SVM, LogR, and RF, and an Ensemble of MNB, LSVC, and RF classifiers, on the labeled dataset. Our main findings suggest that machine learning may be used to classify emails based on their emotional content. However, some models outperform others. During the testing phase, we also discovered that the LogR and LSVC models performed the best, with an accuracy of 72%, while the MNB classifier performed the poorest. Furthermore, classification performance differed depending on whether the dataset was balanced or imbalanced. We conclude that machine learning models that employ emotions for email classification are a promising avenue that should be explored further.
Fish Image Classification using Transfer Learning Method withAdaptive Learning Rate Suhana, Rizka; Mahmudy, Wayan Firdaus; Budi, Agung Setia
Knowledge Engineering and Data Science
Publisher : citeus

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

The diversity of fish species in coral reef ecosystems is one of the indications in determining health in coral reef ecosystems. Many Indonesian Fisheries and Marine Research and Development Agency experts carefully classify fish images. A reliable technique for performing image classification is Convolutional Neural Network (CNN). Transfer learning appears and adopts part of CNN, namely the modified convolution layer. The paper aims to solve the fish classification problem using the pre-trained model of Mobilenet V2. The model has a low computational process and does not use too many memory resources when training image data. The research image data used is 49,281 data of various sizes and 18 types of fish. The image is entered into the transformation process (random rotation, random resize crop, random horizontal flip) on the training and test data to produce varied data. After the transformation process, the image data is entered into the training process using the Mobilenet V2 architecture. Testing the Mobilenet V2 architectural model obtained an accuracy score of 99.54%, which is reliable in classifying fish images.

Page 5 of 12 | Total Record : 117


Filter by Year

2018 2025


Filter By Issues
All Issue