Thèse Architectures Symboliques Vectorielles pour une Interprétation du Génome à Grande Échelle H/F - Doctorat.Gouv.Fr
- CDD
- Doctorat.Gouv.Fr
Les missions du poste
Établissement : Université de Montpellier École doctorale : I2S - Information, Structures, Systèmes Laboratoire de recherche : IGMM - Institut de Génétique Moléculaire de Montpellier Direction de la thèse : Daniele RAIMONDI ORCID 0000000311571899 Début de la thèse : 2026-10-01 Date limite de candidature : 2026-09-30T23:59:59 Motivation : Interpréter le génome¹² consiste à modéliser la relation entre génotype et phénotype, un objectif fondamental de la biologie. Y parvenir pourrait transformer la génétique, la médecine et les biotechnologies agricoles : cela permettrait notamment de développer des traitements adaptés au génome de chaque patient, ouvrant la voie à la médecine de précision¹¹, ou encore de concevoir des cultures mieux adaptées au changement climatique et aux enjeux d'accès à l'alimentation. L'interprétation du génome (Genome Interpretation, GI)¹ est un domaine récent de la bioinformatique² qui cherche à comprendre comment les variations du génome influencent les phénotypes¹²¹¹, qu'ils soient quantitatifs¹³ (par ex. la taille), qualitatifs (par ex. la couleur des yeux), pathologiques¹ (par ex. la maladie de Huntington) ou liés à la réponse aux médicaments¹.
Le grand nombre de variants portés par chaque individu¹ rend cependant difficile l'identification de ceux qui contribuent réellement aux phénotypes. Les premières méthodes de GI¹¹¹² ont longtemps été limitées par le manque de données, d'outils mathématiques et informatiques adaptés, ainsi que par la complexité intrinsèque du problème. Malgré les progrès récents de l'apprentissage profond, l'utilisation des réseaux de neurones reste difficile dans ce contexte en raison du caractère fortement sous-déterminé des cohortes génomiques. Même lorsqu'il est possible de séquencer plusieurs milliers de patients et de témoins, chaque individu peut être décrit par plusieurs millions de variables. À l'inverse, les phénotypes dépendent souvent d'un sous-ensemble relativement restreint du génome, ainsi que d'interactions entre variants et avec l'environnement. Les modèles trop complexes risquent donc de surapprendre et restent souvent difficiles à interpréter.
Ce déséquilibre nécessite des représentations capables de : (1) compresser de très grandes quantités d'information génomique structurée dans des vecteurs de taille fixe et robustes au bruit ; (2) préserver les similarités et structures biologiquement pertinentes, telles que les gènes, haplotypes ou motifs ; (3) exploiter directement ces propriétés biologiques pour la prédiction ; et (4) permettre une interprétation explicite des associations apprises.
Le calcul hyperdimensionnel, également appelé Vector Symbolic Architectures (VSA), possède précisément ces propriétés. Il repose sur des hypervecteurs de taille fixe, robustes au bruit, permettant de représenter et combiner de manière compositionnelle des concepts complexes grâce à des opérations algébriques spécifiques.
L'objectif de ce projet est de développer un cadre hyperdimensionnel pour l'interprétation du génome, reposant sur les VSA afin de représenter les génomes de manière compacte, scalable et interprétable, avec comme application principale la prédiction et la compréhension de maladies humaines cliniquement pertinentes telles que les maladies inflammatoires chroniques de l'intestin (MICI).
Le premier objectif sera de développer des méthodes permettant d'encoder les données de séquençage sous forme d'hypervecteurs multi-échelles, tout en préservant leur sémantique biologique grâce à des opérations compositionnelles telles que le binding et le bundling. Le deuxième objectif sera de construire des modèles prédictifs du risque de maladie et de traits quantitatifs à partir de ces représentations, en combinant pré-entraînement non supervisé et apprentissage supervisé, avec un focus particulier sur les MICI. Enfin, le projet développera des méthodes de décodage interprétables afin d'identifier les contributions de chaque gène et de chaque variant aux phénotypes. À terme, il évaluera les VSA comme nouveau paradigme computationnel pour une analyse du génome scalable, interprétable et biologiquement pertinente. The core focus of D. Raimondi's AI for Genome Interpretation lab at the Institut de Génétique Moléculaire de Montpellier (IGMM) is the development of tailor-made Machine Learning (ML) and Neural Networks (NNs) models for Genome Interpretation (GI)1,2. GI is the umbrella term describing the bioinformatics approaches devoted to understanding the relationship between genotype and phenotype, targeting open biological problems ranging from human clinical genetics to plant biology. My aim is to develop a new paradigm of computational methods to overcome the limitations of current GI algorithms used in quantitative and clinical genetics. To do so, I propose innovative and interpretable tailor-made frameworks of Artificial Intelligence (AI) methods to combine genomics data and contextual knowledge of biological processes, with the goal of modeling how the information encoded in our genome leads to the observed phenotypes. My overarching goal is to enable novel applications of Explainable AI (XAI) in genetics3,4,5, precision medicine, drug discovery6 and agricultural technology7.
Raimondi's lab has an ongoing collaboration with Dr. C. Trottier and X. Bry at IMAG, since they are experts in statistical methods such as Linear Mixed Models (LMMs), which are one very often used in different fields of GI, including quantitative genetics, and plant and cattle breeding programs. In the context of this collaboration, Raimondi, Trottier and Bry are currently co-supervising a post-doc funded by the AISSAI shared initiative between Google and CNRS.
Methodology: This project will develop and evaluate a VSA framework for the disease risk prediction of clinically relevant human disorders, such as Inflammatory Bowel Disease (IBD). We will start by prototyping our models on a simple model organism such as yeast, given the availability of high-resolution whole genome sequencing, and hundreds of phenotypes measured in different environmental conditions. Prototyping on a smaller genome will allow us to rapidly debug the problems we might encounter and test different solutions to optimize our protocol to encode sequencing data into hyperdimensional vectors.
After the prototyping phaset, we will apply our VSA-based Genome Interpretation approach on the disease risk prediction of Inflammatory Bowel Disease (IBD) from human exome sequencing datasets on a dbGAP dataset containing ~4000 cases and controls. Genomic variants from VCF files will be encoded into multi-scale hypervectors (variant, gene, pathway levels) using compositional operators such as binding (Hadamard or circular convolution) and bundling (superposition). Different types of VSA encodings (binary, bipolar, ternary, gaussian) will be tested. We will then train our VSA models for disease risk prediction via supervised learning. Model interpretability will be assessed by unbinding hypervectors to recover per-gene and per-variant contributions, validated against known IBD loci and pathways.
What is new in this project: Using VSA for GI has never been attempted before. This is an innovative project that lies at the frontier between AI and biology. This project will bring four core innovations:
Vector-symbolic genome encoding. Encoding sequencing data into ML-understandable formats, like feature vectors or tensors, is not a trivial task, and it is rarely achieved. Several of my previous publications focus on finding solutions to this challenge. In this project, we will explore the use of VSA for this task. Instead of flat base-by-base encoding, we compress sequencing samples into multi-scale hypervectors (variant gene pathway sample) that holographically encode the semantics of the biological units composing the genome, using the binding, bundling and permutation operations on hypervectors. The properties of VSA ensure the amount of information they can store is not bounded by the size of the samples (i.e. the number of variants or nucleotides in a genome), but just by the number of samples. Moreover, the number N of samples that can be stored in a hypervector with D dimensions conveniently scales as D ~ O(log(N)), making them ideal to embed genomes, that are extremely high dimensional objects (the human genome has ~3*10^9 nucleotides).
The use of VSA for disease risk prediction is unprecedented. The application of VSA to Genome Interpretation is entirely novel. While VSA has been explored in robotics, signal processing, and cognitive computing, it has never been used to model the genotype-phenotype relationship. This project represents the first attempt to encode full genomes as compositional hypervectors for interpretable disease risk prediction, opening a new research direction at the intersection of bioinformatics and hyperdimensional computing.
VSA are a computationally efficient way to do machine learning on large scale data. VSA offers a computationally efficient and scalable alternative to conventional ML for large data. Genomes encoded as fixed-size hypervectors allow constant-time operations independent of genome length, and similarity computations reduce to simple vector algebra. This enables fast, memory-efficient learning and inference on large cohorts, in contrast to deep neural networks that require extensive feature engineering, GPU resources, and costly training cycles.
VSA models are inherently interpretable, and they can be used to extract new biological insights on the molecular determinants of disease. Their compositional structure allows one to unbind and inspect contributions of individual variants, genes, or pathways to phenotype predictions. This enables direct biological insight: researchers can trace which genomic elements drive a prediction, quantify their influence, and discover novel genotype-phenotype associations, transforming predictive models into tools for biological interpretation and hypothesis generation. Together, these innovations make VSA not only an under-studied, compact encoder that could be suitable for GI but a tool for mechanistic interpretation, robust prediction, and scalable cohort analysis.
Le profil recherché
The candidate has been already selected
Compétences requises
- Machine learning