VITAM / CERVO Research Centre, Université Laval
Quebec City, Canada (remote, based in Paris)
Research Collaborator, Contract
September 2025 – Present
Contract- Prototyped a speech-to-image pipeline for schizophrenia research, using self-hosted Whisper, transcript segmentation, and a multimodal model to link segments to images, feeding the lab’s existing analysis framework.
- Benchmarked multiple architectures on 1000 audio files.
- Responsible for the hosting side of the experiments: ran every model myself on my own AMD/ROCm GPU workstation with PyTorch.
- Host a small web server giving the team direct access to the transcription results, speeding up a step that had been done entirely by hand until then.