Research

Trustworthy autonomous systems through evidence, boundaries and verification

My current work asks how increasingly autonomous software and AI systems can remain observable, bounded and accountable when model outputs, agent decisions and downstream actions become operationally consequential.

01

Agentic AI Assurance & Runtime Governance

Evidence-conditioned authority, bounded autonomy, escalation, abstention, human approval and auditable runtime decisions.

Questions

  • What evidence should be required before an AI-generated recommendation becomes an operational action?
  • How should authority decrease when provenance, calibration, context or verification evidence is missing?
  • Where should human approval, abstention and fail-closed behavior enter autonomous workflows?

Related work

02

Multi-Agent Reliability & Failure Containment

Failure propagation, intermediate-output verification, context isolation, containment, recovery and workflow observability.

Questions

  • How do unsupported, stale or contaminated outputs propagate through multi-stage agentic workflows?
  • Which controls can contain failures before downstream stages reuse them as trusted context?
  • How can recovery paths remain explicit, inspectable and reproducible?

Related work

03

Software Engineering for AI Systems

Reproducible evaluation, explicit contracts, testing, traceability, policy-as-code and dependable engineering of AI-assisted systems.

Questions

  • How can assurance requirements be represented as executable and testable system behavior?
  • How can evaluation design avoid leakage, hidden assumptions and misleading aggregate metrics?
  • How should traceability and audit evidence be built into AI-assisted software workflows?

Related work

04

AI-Assisted 5G/6G Systems

AI-assisted network management, fault classification, runtime assurance and bounded operational authority for increasingly autonomous telecom systems.

Questions

  • How should model confidence and operational permission be separated in autonomous network management?
  • What evidence is necessary before network AI outputs are allowed to influence consequential actions?
  • How can telecom automation retain auditability and human escalation as autonomy increases?

Related work