This study proposes a method to identify conserved sequence positions in viral data, with the primary goal of facilitating their translation into proteins. The approach aims to support the early detection of sub-sequences that hold promise as vaccine candidates. The method involves five key steps. First, mutation analysis using the Kimura model: analyzed mutations in the viral data using the Kimura model to provide insights into sequence variations. Second, alignment method selection to align sequences effectively, the progressive alignment approach with the neighbor-joining algorithm was chosen. Third, hybrid algorithm for identifying unchanged sequences: a hybrid algorithm that combines progressive alignment and a logic learning machine (LLM) was employed to identify unchanged sequences. Fourth, determining conserved sequences: based on the longest unchanged sequence, conserved regions within the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) data were identified. Fifth, immuno-bioinformatics analysis for vaccine candidate peptides: using immuno-bioinformatics, protein sequences were analyzed to identify potential vaccine candidates based on B cell epitopes. Experimental validation confirmed the algorithm’s ability to identify candidate proteins and generate vaccine-worthy peptides from conserved protein sequences. The streamlined approach focuses on conserved protein sequences, ensuring efficient and targeted identification of vaccine candidates. By leveraging these conserved regions, contributions can be made to the expedited development of vaccines.
Copyrights © 2026