Commission
Research
All research areas →

Agent Evaluation and Coordination

Generated-Code and GPU-Kernel Correctness

Federated Learning and Privacy

Decentralised Systems and Protocol

Methods
Papers
Collaborate
About Contact Bring a technical claim

Technical guides

Hard operational questions, worked through in detail

Each guide starts with a direct answer, then develops the architecture, the evidence behind it with its status, the limits of that evidence, and the questions teams usually ask next.

01

Governed AI Agents: Authority, Approval, and Evidence

A practical architecture for governing tool-using AI agents with identity, purpose, action tiers, approvals, replayable evidence, and human stop conditions.

For AI platform teams, financial services, security engineers, governance leaders, and operators deploying tool-using agents.

02

Production AI Agent Reliability: State, Side Effects, and Recovery

How to design reliable AI agents using explicit state, durable checkpoints, idempotent effects, replay, evaluation, observability, and operator-owned recovery.

For Agent engineers, platform teams, SREs, architects, and technical leaders moving AI workflows from prototype to production.

03

AI-Generated Code Correctness: Test the Semantics, Not the Vibe

Testing LLM-generated CUDA and Triton kernels with operator-aware oracles, adversarial inputs, reproducible failures, and false-positive controls.

For ML systems engineers, GPU programmers, code-generation researchers, evaluation teams, and engineering leaders adopting coding agents.

04

Multi-Agent Software Engineering Before the Pull Request

How task claims, append-only coordination events, replica convergence, and pre-PR observability expose duplicate and conflicting work among coding agents.

For Developer-tool teams, engineering leaders, coding-agent builders, distributed-systems practitioners, and researchers studying AI-assisted software delivery.

05

Local and Embedded AI Inference Across Language Boundaries

A practical decision framework for local versus hosted inference, in-process model execution, formats, bindings, hardware backends, resource limits, and offline operation.

For AI platform teams, edge and mobile engineers, .NET and systems developers, privacy-sensitive organisations, and architects choosing between hosted and local models.

06

Physical AI and Robotics: From Language Plan to Safe Edge Execution

The system boundary between natural-language robot planning, deterministic validation, capability constraints, edge execution, telemetry, recovery, and warehouse learning benchmarks.

For Robotics, warehouse and logistics operators, manufacturing leaders, industrial-AI teams, edge engineers, and applied-ML researchers.

07

Agent-Compatible Tools: Interfaces for Humans and AI Systems

How to design CLIs, APIs, and MCP tools with discoverable capabilities, typed inputs, stable errors, dry runs, scoped credentials, approvals, and audit records.

For Developer-tool teams, platform engineers, API designers, MCP implementers, security engineers, and organisations exposing operational systems to agents.

08

Blockchain and Agent Security: Control the Key, State, and Effect

Design patterns for validator signing policy, atomic cross-rollup actions, compiler boundaries, and paper-first agent-operated market systems.

For Blockchain infrastructure teams, validator operators, protocol and compiler engineers, fintech architects, applied cryptographers, and agent-security practitioners.

The paper-centred research areas are under Research; reusable evaluation protocols are under Methods.