Agent Harness Engineering
State graphs, routing, tool selection, execution environments, artifact lifecycles, retries, recovery, and explicit control boundaries.
AI Forward Deployed Engineer · Rome, Italy, open to relocation
Agent harness engineering, multi-agent orchestration, evaluation, persistent memory, document intelligence, and deterministic reasoning for enterprise workflows.
I work where software engineering, model behavior, customer workflows, and product judgment meet. At TeamSystem, I lead an AI acceleration team that operates like an internal startup: discover a real need, build the system, validate it with domain experts, and move it into production.
I have built with LangChain since its early 0.x era in 2023, helped take an enterprise legal AI product from beta in 2024 to commercial launch in 2025, and then expanded into LangGraph-based harnesses, persistent memory, Vision LLM pipelines, tool-calling agents, and knowledge-graph reasoning.
My strongest work is not a single prompt or agent. It is the harness that gives models state, tools, memory, quality controls, and a path to improve.
State graphs, routing, tool selection, execution environments, artifact lifecycles, retries, recovery, and explicit control boundaries.
Failure-mode analysis, domain-grounded evaluation criteria, quality gates, structured checks, and reproducible technical artifacts.
Persistent semantic, episodic, factual, and preference memory with classification, versioning, history, retrieval, and governance.
Enterprise RAG, hybrid retrieval, Vision LLM extraction, semantic layers, vector databases, and structured data products.
Computational graphs that turn domain rules and relationships into auditable reasoning paths rather than opaque model guesses.
Working directly with technical users and domain experts, moving fluidly from discovery and prototyping to debugging and production adoption.
Selected enterprise work
Proprietary work is described at a non-confidential architectural level.
A reusable first-party harness for generating and managing complex artifacts across legal, construction, procurement, and accounting workflows.
LangGraph state graph, stateful tool routing, layered QA, quality gates, deterministic verification, artifact lifecycle management, and a code-execution environment synchronized with AWS S3.
The system avoids a proliferation of domain-specific prompts and vertically isolated agents. The same harness can support multiple product domains while preserving explicit controls and verification.
A production RAG and agent platform for legal professionals, developed from the early LangChain 0.x era and scaled over millions of legal documents.
Architecture and primary implementation, beta in 2024, commercial launch in 2025, product roadmap, production operation, and direct iteration with lawyers and compliance professionals.
RAG over 3-4 million vectorized documents, ReAct and tool-calling workflows, failure-mode analysis, evaluation criteria, document management, and compliance automation.
A shared memory capability for persistent cross-session context across products and agents.
Semantic, episodic, factual, and preference memories classified as first-class units, versioned and historicized inside the platform orchestrator.
Designed as a reusable platform capability rather than memory embedded inside a single assistant, with attention to contracts, ownership, retrieval, and evolution.
A centralized ingestion layer that converts complex enterprise documents into reusable structured data products.
Tax and fiscal forms, engineering bills of quantities, delivery notes, invoices, and other documents with complex layouts and domain-specific structure.
Vision LLM extraction and normalization feed agentic querying, retrieval, analytics, and product workflows without rebuilding ingestion per use case.
A reasoning backbone for accounting and fiscal workflows where answers must be explainable, auditable, and reproducible.
The model queries a computational graph of domain entities, rules, and relationships, then constructs the response from the graph-derived reasoning path.
Domain logic remains inspectable and testable, reducing reliance on implicit model knowledge for high-consequence compliance answers.
Agentic systems embedded where professionals already work: enterprise data, Microsoft Word, and engineering-document workflows.
A natural-language data agent over a Databricks semantic layer; a tool-calling Word integration for drafting and revision; and entity extraction and retrieval for construction documents.
Reduce workflow switching. The agent should operate through governed tools and existing product surfaces, not force users into a disconnected chat experience.
Open source
My public work focuses on persistent context, MCP, tool compatibility, model-provider correctness, and production failure modes.
A persistent knowledge-graph memory server for AI coding agents. It tracks goals, constraints, strategies, outcomes, preferences, and semantic links to project code.
kg-mcp with setup toolingAn MCP server for intelligently crawling technical documentation and exposing it through hybrid search and a dynamically extracted ontology.
Fixed request construction for GPT-5-family models by selecting max_completion_tokens across the shared base layer and provider-specific paths.
Stopped Azure providers from mutating caller-owned messages, corrupting user text, and crashing on multimodal message content.
Removed an invalid top-level schema combinator that caused Anthropic-backed MCP sessions to reject the entire tools array.
uniq semantics across virtual filesystemsCentralized parsing of -f, -s, and -w in Mirage's generic uniq implementation, preserving the distinction between an unset option and the literal value 0.
Traced duplicated API cost attribution through cloned session parts and proposed tested fixes for session totals and per-model statistics.
Control flow, state, tools, and verification should be explicit system concerns, not hidden inside ever-larger prompts.
Customer failures must become reproducible cases, measurable criteria, and durable improvements that survive the original incident.
Persistent context needs identity, provenance, versioning, deletion, retrieval policy, and clear ownership, not just a vector-store write.
Probabilistic generation can be bounded by schemas, executable checks, graph-derived rules, and verification layers.
Lawyers, compliance officers, engineers, and customers expose the failure modes that internal AI teams otherwise miss.
Selected public appearances and awards with direct evidence links.
Shared the stage with AWS and Eidosmedia for session ISV202 on practical agent-system implementation with Amazon Bedrock, presenting TeamSystem's Legal AI use cases built with services including Textract and OpenSearch Serverless.
Invited to present TeamSystem's new Legal AI capabilities at the Milan Bar Association (Ordine degli Avvocati di Milano), Italy's largest and most prominent bar association, covering practical adoption, professional productivity, data security, and privacy for the Milan legal community.
A one-to-one conversation with Andrea Cabrini on the practical impact of generative AI and how intelligent systems are entering professional legal workflows.
Panel contributor to the roundtable “Contrattualistica: tra digital transformation, intelligenza artificiale generativa e legal design,” moderated by Andrea Cabrini, Director of Class CNBC.
Part of the seven-person winning team in a 130-participant international challenge backed by AWS, Microsoft, and Google. Built an LLM-first document-intelligence system for procurement workflows, combining conversational search, dynamic structuring, and self-generated extraction prompts. The first-place prize was an award week hosted at Google in San Francisco in 2026.
Innovation prize in 2023, 3rd place at an AWS DevOps hackathon in Tirana in 2024, and 2nd place with the highest-ranked AI project at a 120+ participant conference hackathon.
TeamSystem · Leading a four-engineer fast-track AI team and building reusable platform capabilities.
TeamSystem · Enterprise RAG, autonomous agents, customer evaluation, product architecture, and domain delivery.
TeamSystem · AWS applications, multi-tenant data architectures, early ML systems, ISO 27001, and European engineering teams.
Netlex · PHP, MySQL, JavaScript, and legal software products.
I am based in Rome and open to relocation for high-impact AI engineering and deployed engineering roles.