Applied AI Engineer. I build and evaluate LLM systems in production. Writing about evaluation, interpretability, and things I find interesting.
Side A
Selected work
all projects →
Behavioral evaluation framework for agentic LLM systems
Two-stage search engine for public tenders
Internal MCP server for agent tool composition
AI Voice Assistant
Side B
Writing
all posts →
GenAI assistants in the public sector: measuring safety and trust
Measuring safety and trust in a national-scale PA assistant, from daily red-teaming scenarios to per-tender behavioral evaluation.
AI assistants in the public sector: extending them with reusable tools
Extending production AI assistants with reusable primitives and declarative composition, instead of rewriting a tool every time.
Camilla, the AI assistant taking on Italy's bureaucracy
Architecture and governance choices behind a national-scale chatbot for Italian Public Administration.