Master thesis image

@mastersthesis{anzel2020,
  author  = {An{\v{z}}el, Aleksandar},
  title   = {Determining protein N-glycosylation with machine learning methods},
  school  = {Faculty of Mathematics, University of Belgrade},
  year    = {2020},
  type    = {Master's thesis},
  address = {Studentski trg 16, 11158 Belgrade},
  month   = {2},
  note    = {Available at \url{http://elibrary.matf.bg.ac.rs/bitstream/handle/123456789/5013/AnzelAleksandar.pdf?sequence=1}},
}

Most protein functions are dependent on post-translational modifications (PTMs). One of the most common PTMs in eukaryotes is N-linked glycosylation, which represents the process of attaching an oligosaccharide, sometimes also referred to as glycan, to a protein molecule. Changes in N-linked glycosylation have been associated with various diseases in different organisms. Determining whether a protein will be N-glycosylated or not is the first step of finding an accurate position of an attachment. Wet lab experiments for determining protein N-glycosylation and finding the attachment position are expensive and time-consuming. Several computational tools were created for automatically determining the exact position of N-glycosylation, but there are none that predict the existence of the process. Furthermore, most existing tools are based on human or mammalian proteomes, and very few are using protein data sets of other organisms. It is known that N-glycosylation has distinctive characteristics between different organisms, therefore organism-specific tools are much in need. For this thesis, different machine learning classifiers were developed for determining protein N-glycosylation. Also, different techniques were used to overcome unbalanced data problems that are existent in used data sets.