The problem
Recording clean audio for several people usually means a lapel mic on everyone. We set out to capture each person’s voice separately from a single camera-mounted device.
System overview
- InputCameraVideo feedMicrophone arrayOnboard, multi-channel
- VisionPerson trackingML detection of every human in frame
- AudioBeamformingML-based steering toward each person
- OutputPer-person audioOne stream eachReact front-endLive view & controls
My role
Software
The software runs as separate Python processes. A frame buffer captures camera images with OpenCV, a face-detection process runs an MTCNN network in TensorFlow, and the main application tracks each person with a multi-Kalman filter; they exchange data through shared memory and message queues. The main application talks to the audio process and the React interface over REST, and streams live thumbnails of each person to the interface over a WebSocket.
Hardware
My teammates designed the electronics and enclosure alongside the software. The electronics centre on an Nvidia Jetson TX2, with an IMX185 camera, an array of eight MEMS microphones streamed in over USB through a MiniDSP, a DAC for audio out and a 4.3″ touchscreen.
The enclosure holds the Jetson module on its carrier board, the camera, a custom breakout board, a 1400 mAh LiPo and the touchscreen. The microphone breakouts sit in a curved bar beneath it, and a hot-shoe adapter mounts the unit on a camera.






