Search arXiv⌕ Search

arXiv subjects

Lukas Käll

Publications and source records attributed to Lukas Käll.

3 recordsLinked to original sources

dIon: Fragmentation-Based Invariance for Self-Supervised Learning of Tandem Mass Spectra

We introduce a novel invariance for peptide tandem mass spectrometry data, unlocking self-supervised representation learning that improves de novo sequencing of peptides. This invariance exploits the physical relationship between precursor properties (mass and charge) and fragment-ion evidence, without requiring peptide sequence labels. We introduce dIon, which adapts the DINO framework with two latent prediction tasks, both recovering a clean teacher representation: one from a spectrum mixture, using the precursor as a selection query, and one from a partial spectrum with the precursor withheld. The first associates precursor information with fragment-ion evidence; the second prevents representational collapse onto that information alone. Mechanistic probes support both effects, and ablations show that the full objective performs best. Under identical end-to-end training, dIon initialization improves de novo peptide precision over training from scratch by 5.5 and 8.4 percentage points on the held-out MassIVE-KB and Kingdoms test sets, and by 2.3 and 4.8 percentage points with a larger supervised training corpus. The resulting models surpass fully supervised state-of-the-art de novo sequencing models on the diverse, multi-species Kingdoms corpus under the same greedy-decoding protocol. Without peptide labels, dIon learns strong native peptide-similarity geometry compared with other learned models; with limited peptide-supervised adaptation, it achieves the best retrieval and pair-discrimination performance across all representation benchmarks.

cs.LG↗

A graph-based approach for modification site assignment in proteomics

Background In proteomics, the most probable localizations of post-translational modifications are assessed by localization scores evaluating the likelihood of a given modification to occupy a site on a peptide sequence. When identifying highly modified peptides, localization scores for different modifications can return conflicting results, stacking modifications on the same amino acid. Here, we propose a graph-based approach that assigns modifications to sites in a way that maximizes localization scores while avoiding conflicting assignments. Results The algorithm is implemented as both a standalone Python program and in the compomics-utilities Java library. Our graph-based approach showed the ability to match complex combinations of modifications and acceptor sites, allowing the processing of thousands of peptides in a few seconds. Conclusions Our graph-based approach to modification site assignment allows distributing multiple modifications in a way that maximizes individual localization scores. Having an optimal modification site assignment is important for spectrum annotation and biological interpretation.

q-bio.QM↗

ProHap Explorer: Visualizing Haplotypes in Proteogenomic Datasets

In mass spectrometry-based proteomics, experts usually project data onto a single set of reference sequences, overlooking the influence of common haplotypes (combinations of genetic variants inherited together from a parent). We recently introduced ProHap, a tool for generating customized protein haplotype databases. Here, we present ProHap Explorer, a visualization interface designed to investigate the influence of common haplotypes on the human proteome. It enables users to explore haplotypes, their effects on protein sequences, and the identification of non-canonical peptides in public mass spectrometry datasets. The design builds on well-established representations in biological sequence analysis, ensuring familiarity for domain experts while integrating novel interactive elements tailored to proteogenomic data exploration. User interviews with proteomics experts confirmed the tool's utility, highlighting its ability to reveal whether haplotypes affect proteins of interest. By facilitating the intuitive exploration of proteogenomic variation, ProHap Explorer supports research in personalized medicine and the development of targeted therapies.

q-bio.GN↗