Research
Our ongoing research topics:
Superhuman Perception with RF Signals
TL;DR PanoRadar is a rotating single-chip mmWave radar that brings RF imaging resolution close to LiDAR while staying robust in smoke, fog, and darkness. Its combination of novel signal processing and machine learning enables, for the first time, visual recognition tasks such as semantic segmentation and object detection at radio frequency.
Best Demo Award ACM SRC Grand Finals Second Place
TL;DR CartoRadar achieves RF-based 3D SLAM with accuracy rivaling vision-based approaches, overcoming the sparsity and noise of radar measurements with learning. It provides robust mapping and localization in conditions where optical sensors fail, such as smoke, fog, and darkness.
Best Artifact Award
TL;DR HoloRadar reconstructs 3D geometry beyond the line of sight by exploiting multi-bounce RF reflections: a first stage performs multi-return RF imaging, and a second stage carries out reflection-aware scene reconstruction. The result lets a single mmWave radar perceive around corners.
TL;DR SurfRadar recovers surface properties such as permittivity and roughness from high-resolution mmWave measurements, producing RF-based 3D material maps of indoor scenes. This physical understanding complements geometry for robotics and wireless applications.
mmWave Sensing and Communication
TL;DR Mo³Cap uses lightweight millimeter-wave tags on the body to enable motion capture without cameras. It extends RF sensing toward fine-grained, occlusion-resilient body tracking for robotics and interaction.
TL;DR BlinkWise is a minimalist RF add-on for everyday glasses that tracks blink dynamics at millisecond resolution, with all processing running on an edge microcontroller for private, real-time use. Detailed blink dynamics unlock applications in drowsiness monitoring, workload assessment, and dry-eye disease management.
Acoustic & Audio Intelligence
TL;DR SmartDJ turns high-level, declarative instructions into executable audio edits: an audio language model plans step-by-step operations, and editing models carry them out on stereo audio while preserving spatial cues. Users describe the result they want rather than the steps to get there.
TL;DR AV-Twin builds editable audio-visual digital twins of real spaces with just a commodity smartphone, combining mobile room impulse response capture, visual-assisted acoustic field modeling, and differentiable acoustic rendering. Edits to geometry, materials, and layout update both sound and visuals.
TL;DR VERSA leverages the acoustic reciprocity principle to efficiently learn acoustic fields, enabling resounding: re-rendering how a scene sounds from new source positions. This makes spatial audio capture practical with far fewer measurements.
TL;DR AVR brings physically grounded volume rendering to spatial audio, learning neural impulse response fields that obey acoustic wave propagation. Together with the AcoustiX simulator, it synthesizes realistic, pose-consistent spatial sound at unseen positions.
SpotlightDigital Health
TL;DR BlinkWise is a minimalist RF add-on for everyday glasses that tracks blink dynamics at millisecond resolution, with all processing running on an edge microcontroller for private, real-time use. Detailed blink dynamics unlock applications in drowsiness monitoring, workload assessment, and dry-eye disease management.
TL;DR A case study in fundus oculomics for HbA1c assessment, examining how retinal-imaging AI can screen for cardiovascular risk factors. The work discusses clinically relevant considerations for bringing such models into practice.
TL;DR A contactless wireless sensing system that detects when patients use insulin pens and inhalers and flags administration errors. Deployed unobtrusively in the home, it enables continuous monitoring of medication self-administration.
Immersive Media & Content Creation
TL;DR WaveVerse is a prompt-based framework that generates dynamic indoor 4D worlds and simulates phase-coherent RF signals within them via ray tracing. It provides scalable, realistic RF data for imaging and activity-recognition research when real measurements are scarce.
TL;DR MoScale generates human motion hierarchically from coarse to fine temporal scales with next-scale autoregressive models. Cross-scale and in-scale refinement improve text-to-motion quality, and the model generalizes zero-shot to motion editing and completion tasks.
TL;DR SmartDJ turns high-level, declarative instructions into executable audio edits: an audio language model plans step-by-step operations, and editing models carry them out on stereo audio while preserving spatial cues. Users describe the result they want rather than the steps to get there.
TL;DR AV-Twin builds editable audio-visual digital twins of real spaces with just a commodity smartphone, combining mobile room impulse response capture, visual-assisted acoustic field modeling, and differentiable acoustic rendering. Edits to geometry, materials, and layout update both sound and visuals.