Industry background
Wearables measure accurately, but patients can hardly keep them on over long periods. WiFi and millimeter-wave radar impose almost no extra burden, as the devices are inexpensive and compact, keep working once deployed, and are unaffected by lighting conditions. These advantages make wireless signals the primary sensing channel, with vision complementing them with richer observational details in suitable scenarios.
The research began with data collection. Using commercial devices, the group recorded large amounts of wireless signals in real rooms and annotated them with the true states of the people inside, then trained deep learning models to recognize human activities from the signals. Preliminary results confirmed technical feasibility, as deep learning models could determine from signal changes alone whether a room was occupied and what activity was taking place.
WiFi channel state information (CSI) reflects the smallest human motions, as chest movement and limb motion both leave features in the channel readings. Deep learning models parse these features layer by layer, extending the sensing target from activities to breathing and heartbeat, and forming a contactless vital sign monitoring capability.
Voice is another information channel. Microphones already present in a room capture snoring, coughing, conversation, and footsteps, and speech recognition and audio deep learning models convert these sounds into information about sleep quality and daily states. Voice sensing shares the same deployment advantages as wireless sensing, remaining contactless, unobtrusive, and easy to scale.
Vision provides the most informative observations, covering appearance, posture, motion, and scene. With visual deep learning models, camera frames can be converted into posture and behavior analysis, serving as a reference for the other channels. Privacy requirements define the role of vision in the system, where it works as a complementary channel in specific compliant scenarios and supplies details that wireless channels cannot provide.
Each single channel has its own limitations, which makes multimodal fusion the foundation of the system. With deep learning fusion models, wireless signals, voice, and vision complement one another, keeping perception stable in complex environments. Research results must ultimately return to real environments for testing, and the collaboration with Jiangsu Province Hospital keeps the algorithms sharpened against real ward data and clinical needs.