Marcos Carrasco Sacristán, Óscar Pastor López, Ana León Palacio and Juan Carlos Casamayor Rodenas
The rapid growth of omics resources has led to a highly
fragmented and heterogeneous landscape, giving rise to what we refer to as “genomic chaos”, in which data, information, and knowledge
are represented at different levels of abstraction across genomic, transcriptomic, proteomic, reaction, and pathway domains. This heterogene-
ity, together with the lack of consistent cross-domain standardisation,
complicates semantic interoperability, data integration, and the development of bioinformatics systems. In this context, we present the MOSAIC
model (Multi-Omics Structured and Integrated Conceptual model), an
extensible conceptual and logical reference model that structures core
omics entities, their descriptions, and their relationships into an information architecture. The model is designed to capture how variants, genes,
transcripts, proteins, diseases, reactions, pathways, functional roles, and
chemical context relate across omics domains, providing a coherent basis for representing domain knowledge in a structured way. Rather than
defining a specific physical implementation, the model establishes a reusable
logical foundation for aligning and integrating knowledge from complementary domain resources, standards, and future bioinformatics applica-
tions. By addressing the “genomic chaos” generated by inconsistent representations and fragmented repositories, this work proposes a versioned
and extensible model for organising multi-omics knowledge, whose scope
can evolve as curated resources and scientific evidence introduce relevant
biological distinctions.