Denis Tome'

Dr. Denis Tomé

Senior Research Engineer

My focus is efficient on-device human perception: hand detection, gesture and action recognition, and 3D pose for health and fitness. At Apple I shipped the on-device hand-detection model for a new wearable platform and contributed to products across the Apple Vision Pro lineup, working within latency and power budgets that rule out most off-the-shelf architectures.
At Epic Games I worked at the intersection of computer vision and graphics on digital humans: markerless motion capture, data-driven face and tongue animation, and the authoring tools artists use to build characters. I contributed to the MetaHuman Creator and The Matrix Awakens. The speech-driven tongue animation I co-developed won the best demo award at CVPR 2022 and shipped in UE5.5.
My Ph.D. addressed 3D human pose estimation in monocular and multi-view settings, using semi-supervised and self-supervised methods, including egocentric pose from head-mounted cameras. The central problem was training with very little labeled data and reconciling datasets with incompatible annotations. This work was published at CVPR, ICCV and PAMI.
Across these projects I have designed synthetic datasets and generation pipelines that cover conditions real capture cannot reach. Several remain in production use. My earlier work applied deep learning to pedestrian detection for automotive prototypes and compressed convolutional networks to run on low-power embedded hardware; the pedestrian detection work received a EURASIP best paper award.
I received my Ph.D. from University College London (UCL) under the supervision of Prof. Lourdes Agapito and Prof. Gabriel Brostow as a member of the Vision and Imaging Science Group. My research was funded by the SecondHands European project.

Experience

Sept. 2022 - Present
Apple, Sunnyvale, CA
Senior Research Engineer
Designed and shipped efficient on-device models for hand pose and gesture understanding, optimizing for accuracy under strict power constraints.
Engineered and shipped the on-device hand-detection model for Apple Vision Pro and related spatial computing platforms, achieving a roughly 80% latency reduction over the previous baseline without compromising key metrics.
Built and shipped a real-time on-device model for human perception and analysis (Health & Fitness).
Designed a multi-signal auto-annotation system (VLM semantic reasoning + promptable segmentation).
Dec. 2019 - Sept. 2022
Epic Games, Pittsburgh, PA
Research Scientist, Digital Humans
Co-developed speech-driven tongue and facial animation, winner of the CVPR 2022 Best Demo Award and shipped in UE5.5.
Built digital-human authoring tools and ML models for character creation, contributing to MetaHuman Creator and The Matrix Awakens, and was an original contributor to the UE5 MetaHuman Hair Card Generator, which converts strand-based grooms into real-time hair cards.
Developed graph neural networks for one-shot generation of riggable avatars, used as an internal tool for Fortnite, and skeleton-agnostic ML models for 3D human pose that enable animation retargeting across avatars with different skeletons.
Mar. 2019 - Aug. 2019
Facebook Reality Labs, Pittsburgh, PA
Research Intern, Facebook Reality Labs (FRL)
Researched multi-view egocentric 3D human pose estimation from a collection of cameras for Codec Avatars, Meta's high-fidelity photorealistic avatar project.
Researched 3D mesh reconstruction from images.
May 2018 - Nov. 2018
Oculus Research, Pittsburgh, PA
Research Intern
Researched monocular egocentric 3D human pose estimation from a headset-mounted camera for full-body tracking in Oculus VR headsets, published at ICCV 2019 (xR-EgoPose).
Built a synthetic image generation pipeline targeting hard-to-capture poses to improve model performance, producing the xR-EgoPose high-resolution synthetic dataset.
Aug. 2017 - Nov. 2017
SuperMediaFuture, London, UK
Computer Vision Consultant
Trained 3D human pose estimation models for real-time full-body tracking in consumer AR applications.
The resulting prototype helped the company secure $7.5 million in venture capital funding.
Feb. 2016 - Dec. 2019
University College London (UCL), London, UK
Researcher, Vision and Imaging Science Group
Mentor: Lourdes Agapito
Developed semi- and self-supervised deep learning models for 3D human pose estimation, deployed in end-to-end human–robot interaction demos built with the Karlsruhe Institute of Technology (KIT) as part of the EU-funded SecondHands project.
Teaching assistant for Image Processing (M.Sc.) and Introduction to Programming (B.Sc.) — see Teaching.
Jun. 2015 - Jan. 2016
STMicroelectronics, Milan, Italy
Research Intern
Developed an RPN-style CNN model for pedestrian detection, published in Signal Processing: Image Communication and awarded the EURASIP Best Paper Award.
Compressed CNN detectors to reduce memory and power footprint on low-power embedded hardware, published at IEEE ICCE-Berlin 2016 and demonstrated live on embedded devices.

Education

Feb. 2016 — Feb. 2020
Ph.D. in Computer Vision
University College London, London, UK
Advisor: Prof. Lourdes Agapito, Co-advisor: Prof. Gabriel Brostow
Thesis: More is Better: 3D Human Pose Estimation from Complementary Data Sources
Member of the Vision and Imaging Science Group. Research focused on 3D human pose estimation using semi- and self-supervised deep learning techniques, funded by the SecondHands European project.
Sept. 2013 - Oct. 2015
M.S. in Computer Engineering
Polytechnic University of Milan, Milan, Italy
Master's in machine learning and computer vision. Taught in English.
Final mark: 110 cum Laude/110 (Graduation with Honor)
Exam average score: 29.45/30 (4.0 GPA Equivalent)
Sept. 2010 - Sept. 2013
B.S. in Computer Engineering
Polytechnic University of Milan, Milan, Italy
Taught in Italian.
Final mark: 103/110

Honors and Awards

2022
CVPR'22 Best Demo Award
For Speech Driven Tongue Animation.
Improve realistic animations & surpass the uncanny valley by enabling data-driven tongue animation
2021
EURASIP Best Paper Award for IMAGE COMMUNICATION Journal
For the work Deep Convolutional Neural Networks for pedestrian detection, D.Tomè, F.Monti, L.Baroffio, L.Bondi, M.Tagliasacchi, S.Tubaro, Image Communication, Volume 47, September 2016, Pages 482-489
2017
Nominated for "Outstanding Support for Teaching"
To be nominated candidates must have made a huge impact on the students they teach and/or support; Among 86 nominees from approx. 7100 academic staff members (less than 1.3%)
http://studentsunionucl.org/student-choice-teaching-awards-roll-of-honour-2017

Publications

EPOCH: Jointly Estimating the 3D Pose of Cameras and Humans
Nicola Garau, Giulia Martinelli, Niccolo Bisagno, Denis Tome, Carsten Stoll
T-CAP Workshop at ECCV. 2024.
Project PDF BibTeX
HumMUSS: Human Motion Understanding using State Space Models
Arnab Mondal, Stefano Alletto, Denis Tome
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2024.
Project PDF Poster BibTeX
Speech Driven Tongue Animation
Salvador Medina, Denis Tome, Carsten Stoll, Mark Tiede, Kevin Munhall, Alex Hauptmann, Iain Matthews
IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022.
Project PDF Code Data Poster BibTeX Best Demo Award
SelfPose: 3D Egocentric Pose Estimation from a Headset Mounted Camera
Denis Tome, Thiemo Alldieck, Patrick Peluse, Gerard Pons-Moll, Lourdes Agapito, Hernan Badino, Fernando De la Torre
IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI). 2020.
Project PDF Data BibTeX DOI
xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera
Denis Tome, Patrick Peluse, Lourdes Agapito, Hernan Badino
IEEE/CVF International Conference on Computer Vision (ICCV). 2019.
Project PDF Code Data Slides Poster BibTeX DOI
Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
Denis Tome, Matteo Toso, Chris Russell, Lourdes Agapito
International conference on 3D vision (3DV). 2018.
Project PDF BibTeX DOI
Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image
Denis Tome, Chris Russell, Lourdes Agapito
IEEE Conference on Computer Vision and Pattern Recognition (CVPR). 2017.
Project PDF Code Slides Poster BibTeX DOI
Reduced Memory Region Based Deep Convolutional Neural Network Detection
Denis Tome, Luca Bondi, Luca Baroffio, Stefano Tubaro, Emanuele Plebani, Danilo Pau
IEEE International Conference on Consumer Electronics - Berlin (ICCE-Berlin). 2016.
Project PDF BibTeX DOI
Deep Convolutional Neural Networks for Pedestrian Detection
Denis Tome, Federico Monti, Luca Baroffio, Luca Bondi, Marco Tagliasacchi, Stefano Tubaro
Signal processing: image communication. 2016.
Project PDF Code BibTeX DOI

Teaching

Feb. 2016 - Dec. 2019
Teaching assistant
University College London (UCL), London, UK
COMP0026-A7P-T1, COMP0026-A6U-T1 Image Processing, Instructor: Prof. Lourdes Agapito
Overviewing labs, mentor students in their group projects. Correct assignments and examine students in the final test.
Feb. 2016 - Dec. 2019
Teaching assistant
University College London (UCL), London, UK
COMP211P Introduction to Programming, Instructor: Prof. Rae Harbird
Overviewing labs, mentor students in their group projects. Correct assignments.

References

Leon Park, ML Engineering Manager
Apple
Email: jeongwook_park@apple.com
Francisco Vicente, Lead Research Engineer
Carnegie Mellon University (CMU)
Email: Francisco@fkvicente.com
Dr. Hernan Badino, Research Scientist
Reality Labs
Facebook (Meta)
Email: hernan.badino@fb.com
Prof. Fernando De la Torre, Research Associate Professor
Robotics Institute
Carnegie Mellon University (CMU)
Email: ftorre@cs.cmu.edu