Claim Missing Document
Check
Articles

Found 2 Documents
Search

IndoBERTSkill: pretrained domain-specific language model for recognition Indonesian skill Meilany Nonsi Tentua; Suprapto Suprapto; Afiahayati Afiahayati
International Journal of Advances in Intelligent Informatics Vol 12, No 2 (2026): May 2026
Publisher : Universitas Ahmad Dahlan

Show Abstract | Download Original | Original Source | Check in Google Scholar

Abstract

The pretrained language model in Indonesian is already available for natural language processing tasks. However, this pre-trained model has been trained on Indonesian text, which has a different structure from the job description. Due to this, the pre-trained language model effectiveness for skill recognition purposes. IndoBERTSkill is a novel pre trained domain-specific language model that recognizes Indonesian language skills. It is built on the Bidirectional Encoder Representations from Transformers (BERT) architecture. IndoBERTSkill was trained on an extensive collection of Indonesian language texts from the Indonesian Wikipedia, the English Wikipedia, and the Indonesian Job Description from the job portal. IndoBERTSkill's performance was evaluated through two main approaches: (1) language modeling via Masked Language Model (MLM) prediction, and (2) fine-tuning on a custom annotated dataset (NERSkill) for Named Entity Recognition (NER) tasks. The fine-tuning process involved training a classification layer on top of the IndoBERTSkill model using BIO tagging to identify hard skills, soft skills, and technology entities. Similarly, the skill recognition model derived from IndoBERTSkill exhibits the highest F1-Score among various pre-trained language models, precisely at 87%, thus demonstrating robustness and strong generalizability for skill entity recognition in Indonesian job descriptions. IndoBERTSkill provides valuable resources for developing Indonesian natural language processing applications that require skills introduction. This could increase the accuracy and efficiency of skills recognition across various domains, including job matching, education, and training.
Classification of Nicotine Treatment Response Based on Gene Expression Profiles Using Support Vector Machine and Gaussian Process Models Rahmadi Yotenka; Adhitya Ronnie Effendie; Gunardi Gunardi; Afiahayati Afiahayati; Aisha Ellany Midnova; Eugenia Rivanda Gita Flamboyan; Daninta Indiana Mahaputri; Leonardo Leonardo; Nayla Revania Dewayani; Muhammad Ahnaf Billie Chesta
Mathematical Journal of Modelling and Forecasting Vol. 4 No. 1 (2026): June 2026
Publisher : Universitas Negeri Padang

Show Abstract | Download Original | Original Source | Check in Google Scholar | DOI: 10.24036/mjmf.v4i1.50

Abstract

Nicotine is known to impair endothelial function and increase cardiovascular risk through transcriptional dysregulation. This study investigates the gene expression response of human induced pluripotent stem cell (iPSC)-derived endothelial cells to nicotine exposure using the RNA-seq dataset GSE274506. The analysis was conducted on 40 samples, consisting of 20 nicotine-treated and 20 untreated/control samples, using a 5-fold outer stratified cross-validation with 3-fold inner cross-validation for model tuning, reflecting the limited sample size relative to the high-dimensional gene expression feature space. Differential expression analysis identified 46 significant genes, comprising 28 upregulated and 18 downregulated, indicating perturbations in G protein-coupled receptor (GPCR) signaling, calcium homeostasis, and inflammatory processes. Functional enrichment analyses based on Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), and Reactome consistently revealed dominant involvement of GPCR signaling, cyclic adenosine monophosphate (cAMP) signaling, calcium signaling, and transient receptor potential (TRP)-related pathways, suggesting a coordinated molecular response to nicotine-induced stress. To discriminate between nicotine-treated and control samples, Support Vector Machine (SVM) and Gaussian Process Classification (GPC) models were evaluated. The linear SVM achieved the best and most stable performance, with an accuracy of 0.875, an F1-score of 0.881, and a G-mean of 0.861, outperforming SVM with radial basis function kernels, single-kernel GPC variants, and a multiple kernel learning (MKL) GPC model. These findings indicate that the underlying transcriptomic structure of the data is predominantly linear, favoring linear kernel-based classifiers in high-dimensional gene expression analysis.