Our research — published.

Peer-reviewed and preprint publications from the Emobot team and our academic partners. Every headline stat you see on this site traces back to one of these papers.

7 publicationsASCP 2026 · IEEE ICASSP · Springer LNCS · ASSTA · arXiv · medRxivCentraleSupélec · IETR · Emobot

Clinical validation

2 papers

Clinical validationMarch 8, 2026

Clinical Validation of the EMOCARE-Derived Depressive Symptom Severity Score using Established Clinician- and Self-reported Scales: Preliminary Evidence Across 3 Prospective Studies

Authors: Emobot clinical team et al.

Venue: medRxiv preprint · JMIR Mental Health (in review)

Pooled analysis of three prospective studies in adults with mood disorders, comparing Emobot's passively-derived multimodal depression score against the clinician-rated MADRS and patient-reported PHQ-9. Reports moderate-to-strong convergent validity and sensitivity to symptom change — the headline MADRS r = 0.89 / PHQ-9 r = 0.83 figures quoted across Emobot materials trace back to this dataset.

  • MADRS repeated-measures correlation r ≈ 0.895 (clinician gold standard)
  • PHQ-9 delta correlation ρ ≈ 0.834 (patient self-report)
  • Sensitivity to change preserved across three independent cohorts
  • Anchors Emobot's clinical-validation narrative across US/EU deployments

DOI: 10.64898/2026.03.08.26347894v1

Clinical validationNovember 2024

An Alternative Approach to Depression Diagnosis: Predicting Individual Depressive Symptoms from Speech

Authors: Karim M. Ibrahim · Antony Perzo et al.

Venue: SST 2024 — 20th Australasian International Conference on Speech Science and Technology (ASSTA), pp. 167–171

Explores whether acoustic features extracted from short patient speech excerpts can predict individual depressive symptoms — not a single overall depression score, but symptom-level signal (e.g. anhedonia, psychomotor slowing). Moves the field from "depressed vs. not" classification toward the symptom-level granularity clinicians actually use.

  • Predicts individual PHQ-9-style symptoms from speech, not just an overall label
  • Uses valence / arousal acoustic features derived from vocal biomarkers
  • Evidence for continuous, symptom-level monitoring between clinic visits
  • Clinically motivated: matches how interventional psychiatrists actually titrate care

Conference proceedings

1 paper

Conference proceedingsMay 27, 2026

Continuous Multimodal Passive Monitoring of Depressive Symptoms via Smartphone — Clinical Validation of EMOCARE

Authors: Antony Perzo · Michael Todd Sapko, M.D., Ph.D. · Tanel Petelot · Renaud Séguier · Jonathan C. Javitt, M.D., M.P.H.

Venue: ASCP 2026 · Late-Breaking Poster · Annual Meeting of the American Society of Clinical Psychopharmacology · Emobot × NRx Pharmaceuticals × CentraleSupélec × Johns Hopkins

First public pooled clinical validation of EMOCARE, Emobot's passive multimodal Depression Thermometer. Interim analysis across three prospective observational studies (EMC1, EMC2-FR, EMC2-BD) in n = 45 adults with Major Depressive Disorder or Bipolar Disorder compared the continuous, zero-burden 0–100 EMOCARE score — derived on-device from facial, vocal, actigraphy and digital-behavior signals — against the clinician-rated MADRS and HAM-D₁₇ and the self-reported PHQ-9 and GAD-7. Co-authored with NRx Pharmaceuticals (Nasdaq: NRXP) and presented as a late-breaking poster at ASCP 2026 by Tanel Petelot.

  • Within-person concordance vs. clinician-rated MADRS: r = 0.895 (repeated-measures correlation, p = .016)
  • Sensitivity to symptom change vs. PHQ-9: ρ = 0.834 (consecutive-visit Δ, p < .001)
  • Concurrent association across MADRS, HAM-D₁₇, PHQ-9 and GAD-7: ρ = 0.61–0.83
  • Feasibility at scale: continuous EMOCARE scores generated across the full follow-up in the pooled cohort, ≥7 valid days / 14-day window data-density criterion confirmed
  • On-device feature extraction — raw images and audio never leave the patient's phone

Speech emotion AI

2 papers

Speech emotion AIAugust 19, 2025

EmoSLLM: Parameter-Efficient Adaptation of LLMs for Speech Emotion Recognition

Authors: Hugo Thimonier · Antony Perzo · Renaud Seguier

Venue: arXiv preprint (2508.14130) · CentraleSupélec × Emobot

Proposes a parameter-efficient way to turn a general Large Language Model into a speech emotion recognizer. Audio features are extracted, mapped into the LLM's representation space via a learnable interface, then fused with the transcript and a task prompt. Low-Rank Adaptation (LoRA) keeps the fine-tuning footprint small.

  • Outperforms nearly all existing Speech-Text LLMs on standard SER benchmarks
  • Uses less than half the trainable parameters of competing approaches
  • Couples acoustic and linguistic cues — the combination that matters for clinical-grade emotion signal
  • Directly feeds Emobot's vocal-biomarker pipeline

arXiv: 2508.14130

Speech emotion AIApril 2024

Towards Improving Speech Emotion Recognition Using Synthetic Data Augmentation from Emotion Conversion

Authors: Karim M. Ibrahim · Antony Perzo · Simon Leglaive

Venue: IEEE ICASSP 2024 — International Conference on Acoustics, Speech and Signal Processing, Seoul

Speech Emotion Recognition is starved for labelled data. This paper trains an end-to-end speech-to-speech emotion conversion model (HuBERT → unit-translation → HiFi-GAN) to generate synthetic expressive audio, then uses it to augment training for a wav2vec 2.0 emotion classifier. Evaluates both perceptual quality (MOS) and downstream SER accuracy on IEMOCAP and RAVDESS.

  • IEMOCAP speaker-dependent accuracy: 76.19% with original + synthetic vs 74.16% on original alone
  • IEMOCAP speaker-independent: +2.09 pp over the original-only baseline
  • RAVDESS speaker-dependent: 93.05% accuracy (+1.97 pp)
  • Optimal augmentation ratio around 0.75× synthetic:original across most settings
  • Opens the door to less-resourced languages where labelled emotional speech does not exist

DOI: 10.1109/ICASSP48485.2024.10445740

Representation learning

2 papers

Representation learning2025

Can AI Decode the Circumplex Model of Affect? A Data-driven Study

Authors: Amdjed Belaref · Samir Sadok · Karim M. Ibrahim · Zineb Noumir · Renaud Seguier

Venue: ICPR 2024 Workshop · Springer Nature LNCS 15614, pp. 97–108 (2025)

Tests whether the latent spaces learned by modern Transformer models reproduce Russell's circumplex model of affect — the psychological framework that organises emotions in a valence × arousal circle. Applies dimensionality reduction to text-only, audio-only, and multimodal representations, then measures how faithfully each reproduces Russell's ordering of eight affective words.

  • Unimodal (text-only or audio-only) models only partially recover Russell's circle
  • Multimodal (text + audio) representations closely replicate the circular structure
  • Provides a data-driven validation of the valence-arousal model Emobot uses clinically
  • Strengthens the case for multimodal fusion over any single-signal approach

DOI: 10.1007/978-3-031-87657-8_7

Representation learning2024

T-JEPA: Augmentation-Free Self-Supervised Learning for Tabular Data

Authors: Hugo Thimonier et al.

Venue: OpenReview (peer-reviewed conference submission)

Adapts the Joint Embedding Predictive Architecture (JEPA, from LeCun's world-models line of work) to tabular data — predicting representations rather than raw values, and doing so without hand-crafted data augmentations. Foundational methodology that underpins how Emobot's research stack learns compact, transferable representations from heterogeneous patient signals.

  • Augmentation-free self-supervised pre-training on tabular data
  • Competitive with augmentation-based baselines without domain-specific tricks
  • Transfers to downstream tabular classification and regression tasks
  • Methodological foundation for learning from multimodal clinical features

More papers in the pipeline

The EMC1, EMC2-FR, EMC2-BD, EMC1-MDD-US, and REMOOD studies are producing manuscripts as results are finalised. This page is updated as new publications clear peer review or land on the preprint servers.

Want to see how the method works in a real clinic?

A 30-minute demo walks through the dashboard, the patient app, and the validation dataset that sits behind the headline numbers.