The Phase 1 corpus is complete, annotated end to end, and available for enterprise licensing today.
studio recordings
of audio, about 9.6 hours
multi-stem masters
vocals, sitar, sarangi, bansuri, harmonium, tabla
Hindustani classical / Bollywood
of standardized metadata per track
consent-logged with a per-track royalty trail
compliant: Findable, Accessible, Interoperable, Reusable
Phase 2 expands the corpus toward 500+ hours and adds the Carnatic tradition. Enterprise partners get priority access.
Each track carries a metadata record designed for both training pipelines and musicological analysis, organized across five categories:
Raga classification, tala identification, melodic markers, tempo, dynamics.
Artist lineage (gharana), experience, training history.
Raga etymology, emotional framework (rasa theory), seasonal appropriateness.
Pitch contour data, ornamentation catalogue (gamakas), rhythmic complexity metrics.
Artist compensation records, consent documentation, archival status.
Six tracks, one per instrument, drawn directly from the corpus. Each preview streams at reduced resolution with an audible watermark, and each card shows a live excerpt of the track's annotation JSON. What you hear is a demo. What we license is 96kHz / 24-bit multi-stem masters with the complete 80-field record.
Rarest instrument in Western datasets
Get all six tracks as full-length 96kHz / 24-bit WAV stems with the complete 80-field JSON for each, delivered under a two-week evaluation license.