Skip to main content
NJUPT emblem
Back to research directions

Research Direction 01

Biomedical Multimodal Sensing: Sensing Life in Silence

Machine learning and deep learning models extract physiological and behavioral information from wireless signals, and a health assistant agent built on large language models (LLMs) integrates and presents the results, all without wearables, toward real clinical deployment.

Direction Lead

Xuwen Zhang photo

Direction Lead

Xuwen ZhangUndergraduate Student

barcaxu@outlook.com

Research Partners

  • Jiangsu Province Hospital logoJiangsu Province Hospital
  • FutureComm Lab, SEU logoFutureComm Lab, SEU

Research Direction

Multimodal Sensing

WiFi and CSI signals travel through wards and living rooms at all times, and the breathing, movements, and posture changes of the human body leave measurable traces in these signals. Microphones capture snoring, coughing, and conversation, while cameras record posture, motion, and scene information in specific compliant scenarios. At the data processing layer, machine learning and deep learning models extract features from the data of each channel and map them to localization, activity, vital signs, posture, and behavior, enabling long-term contactless detection of arrhythmias such as atrial fibrillation. Multimodal fusion is performed by deep learning models, which integrate the results of all channels into a unified perception outcome. At the reporting layer, the group has built an integrated personal health assistant agent on large language models (LLMs), which organizes observations, explains model judgments, and supports screening, diagnosis, and follow-up. The collaboration with Jiangsu Province Hospital validates this complete technical pipeline in real clinical settings.

Core tasks

In collaboration with Jiangsu Province Hospital, the group collects and annotates multichannel sensing data in real wards, trains and validates machine learning and deep learning models, studies multimodal fusion, and builds a health assistant agent on large language models (LLMs), so that human localization, physiological and behavioral sensing, human pose estimation, and Alzheimer's disease screening hold up to real environments.

Industry background

Wearables measure accurately, but patients can hardly keep them on over long periods. WiFi and millimeter-wave radar impose almost no extra burden, as the devices are inexpensive and compact, keep working once deployed, and are unaffected by lighting conditions. These advantages make wireless signals the primary sensing channel, with vision complementing them with richer observational details in suitable scenarios.

The research began with data collection. Using commercial devices, the group recorded large amounts of wireless signals in real rooms and annotated them with the true states of the people inside, then trained deep learning models to recognize human activities from the signals. Preliminary results confirmed technical feasibility, as deep learning models could determine from signal changes alone whether a room was occupied and what activity was taking place.

WiFi channel state information (CSI) reflects the smallest human motions, as chest movement and limb motion both leave features in the channel readings. Deep learning models parse these features layer by layer, extending the sensing target from activities to breathing and heartbeat, and forming a contactless vital sign monitoring capability.

Voice is another information channel. Microphones already present in a room capture snoring, coughing, conversation, and footsteps, and speech recognition and audio deep learning models convert these sounds into information about sleep quality and daily states. Voice sensing shares the same deployment advantages as wireless sensing, remaining contactless, unobtrusive, and easy to scale.

Vision provides the most informative observations, covering appearance, posture, motion, and scene. With visual deep learning models, camera frames can be converted into posture and behavior analysis, serving as a reference for the other channels. Privacy requirements define the role of vision in the system, where it works as a complementary channel in specific compliant scenarios and supplies details that wireless channels cannot provide.

Each single channel has its own limitations, which makes multimodal fusion the foundation of the system. With deep learning fusion models, wireless signals, voice, and vision complement one another, keeping perception stable in complex environments. Research results must ultimately return to real environments for testing, and the collaboration with Jiangsu Province Hospital keeps the algorithms sharpened against real ward data and clinical needs.

Research directions

One set of techniques, four directions. Starting from data collected in real environments, the group makes its models deliver usable answers in human localization, physiological and behavioral sensing, human pose estimation, and Alzheimer's disease screening.

Human localization

Indoor position and trajectory are the most basic sensing information. Machine learning and deep learning models analyze feature changes in WiFi signals and radar echoes to determine the locations of patients and medical staff, and to detect anomalies such as a fall followed by prolonged immobility.

Human localization is the foundation of smart wards, supporting patient flow management, fall alerts, and path planning for medical robots. Models are evaluated across rooms and layouts, so localization remains stable after the environment changes.

Physiological and behavioral sensing

Vital signs are reflected in the weakest signal variations. After CSI and radar signals are collected and preprocessed, machine learning and deep learning models learn the relationship between these variations and physiological signals, enabling continuous estimation of respiration, heart rhythm, and sleep apnea events without electrodes or skin contact.

The same models can identify persistent and intermittent rhythm abnormalities from long-term observations, including atrial fibrillation detection. Data from real wards and the collaboration with Jiangsu Province Hospital are used for training and validation, connecting signal-level modeling with clinical monitoring needs.

Human pose estimation

Deep learning models extract spatial and temporal features from radar point clouds and signal changes to reconstruct human posture, including sitting, lying, walking, and falling. The models learn body structure and motion patterns from data, so results remain available at night and under partial occlusion without depending on lighting.

The estimated pose is then converted into measurable clinical indicators, including fall events, post-operative mobility, and motion quantification during rehabilitation. These indicators provide continuous data for assessment, while multimodal fusion can compare pose changes with physiological and behavioral signals.

Alzheimer's disease screening

Alzheimer's disease leaves characteristic marks in daily behavior years before diagnosis, such as frequent rising at night, disrupted circadian rhythms, slowing gait, and declining activity levels. These changes are subtle and gradual, difficult for family members to notice, and clinic consultations rely on the recollections of patients and their families.

WiFi, CSI, voice, and visual data provide complementary evidence for these changes, and machine learning and deep learning models identify gradual alterations in sleep, movement, speech, and daily activities.

Long-term multimodal data can support early screening and condition follow-up, capturing changes that are difficult to detect in a single clinic visit. Large language models organize these observations into understandable reports and explain them through health assistant conversations, while clinical decisions are grounded in model evidence, physician review, and the collaboration with Jiangsu Province Hospital.

Future directions

In the ward of the future, sensing models will run on edge devices at the bedside, watching the room around the clock. Shallow breathing, abnormal heart rhythm, leaving the bed at night, and accidental falls will all be recognized in time and reported to the nurses' station through the health assistant agent. Patients will not need to wear any device, monitoring is performed by the environment itself, and continuously recorded physiological data will provide doctors with a long-term basis for diagnosis. This is already becoming reality, as the group's self-developed integrated millimeter-wave, WiFi, and infrared acquisition platform can already capture respiration, heart rhythm, and body movement synchronously in real rooms, and the contactless atrial fibrillation monitoring project in collaboration with Jiangsu Province Hospital is accumulating clinical data, so that the detection of paroxysmal rhythm anomalies no longer depends on chance.

The same capability will leave the hospital and reach millions of households. A router and a smart speaker together form a home sensing system that keeps recording sleep, breathing, and activity rhythms over long periods. The health assistant agent tracks the slow changes in health hidden within data spanning dozens of days, bringing early screening into everyday life. Silent rhythm anomalies such as atrial fibrillation, as well as the earliest signs of Alzheimer's disease, will be flagged in time, with the evidence explained and suggestions offered through conversation. Post-discharge reviews and follow-ups will also be completed at home, where recovery progress is continuously recorded and medical attention is suggested only when something abnormal appears, so routine checks no longer require repeated trips to the hospital. Every elderly person's home will have a tireless personal doctor, and long-term observation from the everyday environment will replace judgments formed in a single clinic visit.

The continuous observations gathered from wards and homes eventually point to the same question, how to see future changes in the data we already have. World models are the group's answer. A world model learns from long-term multimodal observations how human states evolve, predicting the course of recovery and flagging the warning signals before events such as rhythm anomalies, so the care team is prepared before change arrives. Beyond prediction, world models support interaction, where physicians can ask about likely outcomes under different diagnosis and treatment plans, and families can explore the possibilities in natural language, with every answer grounded in simulated trajectories and observable data.

Behind all of this lies the group's ongoing research on unified representations and multimodal large models, which allow models to transfer across scenes and allow sensing systems to explain their own judgments and converse naturally with people, and world models stand on exactly this foundation. As perception spreads into every ward and every household, screening, diagnosis, and follow-up will form one unbroken chain, where diseases are observed before they are diagnosed and risks are flagged before they occur, and healthcare will move from event-driven to always present. And it all begins with one model and one dataset...

Representative works

Millimeter-Wave, WiFi, and Infrared Integrated Real-Time Acquisition Platform

Experimental platform · Multimodal sensing

The group has built an integrated real-time acquisition platform combining millimeter-wave radar, WiFi, and infrared sensing, supporting synchronized multi-source data collection and algorithm validation for vital sign monitoring and complex scene experiments.

Millimeter-wave, WiFi, and infrared integrated real-time acquisition platform
Integrated real-time acquisition platform

AceNet: attention-guided context enhancement for imbalanced action recognition via RF signals

Authors: Biyun Sheng, Hao Liu, Hui Cai, Yiping Zuo, Jian Zhou, Fu Xiao

IEEE Transactions on Mobile Computing · 2026

This work addresses class imbalance in activity recognition with RF signals and proposes an attention-guided context enhancement network to distinguish semantically similar actions and reduce decision boundary bias caused by imbalance.

View paper

CoSense: Respiratory Detection Based on WIFI CSI and Millimeter-Wave Radar

Authors: Jiaming Liu, Xiaoxiao Qiao, Zeyuan Wu, Weibei Fan, Yiping Zuo, Xin He

2026 29th International Conference on Computer Supported Cooperative Work in Design (CSCWD) · 2026

This work proposes a contactless respiration detection system based on joint WiFi CSI and millimeter-wave sensing, fusing signal features to improve respiratory detection accuracy in complex indoor environments.

View paper

Adaptive Hybrid Routing for Wi-Fi CSI-Based Indoor Human Activity Recognition Using DTW-KNN and SVM

Authors: Jiasheng Song, Yiping Zuo, Chen Dai, Weicong Chen, Ning Gao, Xin He, Weibei Fan

2026 29th International Conference on Computer Supported Cooperative Work in Design (CSCWD) · 2026

This work studies WiFi CSI-based indoor human activity recognition and proposes an adaptive hybrid routing method combining DTW-KNN and SVM for robust recognition in resource-constrained scenarios.

View paper