PD Voice Research
Research Project · Deep Learning · Voice Analysis

Voice-Based Detection of Parkinson's Disease Using Deep Learning

Exploring non-invasive early screening through acoustic biomarkers in sustained vowel phonation — bridging deep learning and clinical neurology.

Research demo — not a clinical toolCNN-LSTM hybrid architectureMel-spectrogram analysis

The Challenge at Scale

Why This Research Matters

10M+
People Affected
worldwide and rising with aging populations
2nd
Most Common
neurodegenerative disease after Alzheimer's
90%
Voice Changes
of PD patients experience vocal symptoms
87%+
Model Accuracy
on subject-independent cross-validation

Methodology

How the Analysis Works

From a simple voice recording to a deep learning assessment — three stages transform raw audio into acoustic feature representations.

01

Voice Recording

Participants sustain a vowel sound ('ahh') for several seconds. These recordings capture subtle neuromuscular tremors and instabilities invisible to the naked eye.

02

Signal Processing

Audio is converted into mel-spectrograms — 2D time-frequency images that reveal acoustic biomarkers across the phonation window.

03

Deep Learning

A hybrid CNN-LSTM architecture extracts spectral features and models temporal dynamics, producing a probabilistic score for Parkinson's indicators.

Research Summary

A Non-Invasive Path to Early Detection

Current Parkinson's diagnosis relies on observing motor symptoms, which often appear only after significant neurological damage — sometimes years after disease onset. This limits the effectiveness of neuroprotective interventions.

Voice changes, however, can appear much earlier. This research investigates whether a hybrid CNN-LSTM deep learning model, trained on mel-spectrogram representations, can distinguish between PD patients and healthy controls in a subject-independent evaluation framework.

The model captures both local spectral patterns and long-range temporal dynamics — the subtle tremor and aperiodicity in vocal production characteristic of Parkinson's dysphonia.

CNN-LSTM
Architecture
Mel Spectrogram
Input
Subject-Independent
Evaluation
87.3%
Accuracy

Explore the Research

Ready to See It in Action?

Dive into the full methodology, model architecture, and results — or try the interactive demo to see the pipeline in action.