SEA: Shared Embeddings for Astrophysics

This is an ongoing project at NCRA-TIFR, supervised by Dr. Yogesh Wadadekar in collaboration with Saptadip Saha, focused on learning shared embeddings across different astronomical data modalities.

Galaxies are studied through many different modalities - images, spectra, and morphological descriptions - but these are usually analyzed in isolation. Learning a shared representation across them would make it possible to connect and reason about the same galaxy’s information jointly, and to retrieve or compare galaxies across modalities.

Traditionally, different modalities have been studied using separate models, making it difficult to combine information across them. Thus, learning a shared embedding space would be valuable, and contrastive representation learning offers a way to do so, aligning different modalities of the same object directly from data.

The goal is to build SEA (Shared Embeddings for Astrophysics), a model that learns a common embedding space across galaxy images, their morphological text descriptions, and spectral data using contrastive representation learning.

SEA pipeline: galaxy images, morphological text, and spectra encoded into a shared embedding space
Figure 1: SEA pipeline - images, text, and spectra encoded into a shared embedding space.

Status: in progress - full write-up, code, and model weights to follow once results are public.