← Work / Projects

Team & software lead · Capstone · 2019

Dynamic Shotgun Microphone

A camera-guided microphone that records a separate audio stream for every person in the shot, with no lapel mics. Computer vision finds each person; ML-based beamforming steers a microphone array to them in real time.

Role
Team lead · software lead
Team
5 engineers
When
Sep 2018 – Apr 2019, final year at UWaterloo
Stack
Python on an Nvidia Jetson TX2 · real-time face & head tracking · React front-end
Prototype microphone array and camera mounted on a DSLR
Prototype mounted on a DSLR
1 stream
Per person in frame, captured without lapel mics
Real time
Video tracking and beam steering
5 · 6 mo
Team size and build time

The problem

Recording clean audio for several people usually means a lapel mic on everyone. We set out to capture each person’s voice separately from a single camera-mounted device.

System overview

  1. Input
    CameraVideo feed
    Microphone arrayOnboard, multi-channel
  2. Vision
    Person trackingML detection of every human in frame
  3. Audio
    BeamformingML-based steering toward each person
  4. Output
    Per-person audioOne stream each
    React front-endLive view & controls
Camera-guided beamformingSimplified

My role

As team and software lead I defined the product scope from our timeline and customer interviews, and wrote the Python backend that processes video in real time and talks to the React front-end and the audio stack.

Thumbnail of the demo video: the app’s interface, with a round thumbnail of each person’s face above a live camera view of the two of them
Demo video

Software

The software runs as separate Python processes. A frame buffer captures camera images with OpenCV, a face-detection process runs an MTCNN network in TensorFlow, and the main application tracks each person with a multi-Kalman filter; they exchange data through shared memory and message queues. The main application talks to the audio process and the React interface over REST, and streams live thumbnails of each person to the interface over a WebSocket.

Software architecture: a React user interface in Chrome, and Python processes for the main application, audio, frame buffer and face detection, linked by REST APIs, a WebSocket and message queues
Software architecture

Hardware

My teammates designed the electronics and enclosure alongside the software. The electronics centre on an Nvidia Jetson TX2, with an IMX185 camera, an array of eight MEMS microphones streamed in over USB through a MiniDSP, a DAC for audio out and a 4.3″ touchscreen.

Electrical block diagram: the Jetson TX2 with its camera, display, audio, network and storage interfaces, connected to the IMX185 camera, eight-microphone array, MiniDSP USB streamer, DAC, touchscreen, battery gas gauge, USB-C power controller and IMU
Electrical architecture

The enclosure holds the Jetson module on its carrier board, the camera, a custom breakout board, a 1400 mAh LiPo and the touchscreen. The microphone breakouts sit in a curved bar beneath it, and a hot-shoe adapter mounts the unit on a camera.

Exploded CAD view of the enclosure with 11 numbered parts, including the MiniDSP, DAC, IMX185 camera, Jetson TX2 module and carrier board, custom breakout board, touchscreen, LiPo battery, MEMS microphone breakouts and hot-shoe adapter
Exploded view of the enclosure

Related