Real-time ETL: MongoDB → BigQuery
Python multiprocessing pipeline with automatic schema detection from the MongoDB OpLog, containerised for deployment.
Data & AI Engineer
I build LLM-powered systems and the data platforms behind them — from AI agents and semantic search to real-time streaming pipelines and high-performance scientific simulations.
I am a Mathematical Engineer with a master's in Computational Science and Computational Learning from Politecnico di Milano (110/110 cum laude). My work sits at the intersection of AI & Data Science, Data Engineering and Scientific Computing, with a strong focus on designing and deploying LLM-powered systems and scalable data platforms.
My experience spans production pipelines for AI agents, document extraction and semantic search, alongside big data, cloud data warehousing and real-time processing — from high-throughput Spark and Kafka streaming to cardiac-muscle modelling on HPC clusters.
Since 2018 I am a Co-Founder and Data Advisor at thefaculty, where I help shape data and AI strategy for product teams.
Building modern data platforms with FastAPI services and cloud data engineering on AWS.
Shaping data and AI strategy for product teams, after leading the data function as Head of Data.
Java/Spring Boot microservices and Scala/Spark streaming on AWS EMR, with high-throughput Spark and Kafka pipelines (1000+ msg/sec) for electric-grid monitoring.
Python multiprocessing pipeline with automatic schema detection from the MongoDB OpLog, containerised for deployment.
Finite-element modelling in C++ with simulations run on an HPC cluster.
Transfer learning and ensemble methods, reaching 95% test accuracy (top 20% of the leaderboard).