dst9.tech Applied AI Agent Evaluation

Independent international research group

Evaluating AI Agents in Isolated Real-World Environments

We operate partner-specific evaluation environments where AI agents can be tested in controlled professional workflows, with defined access boundaries, human oversight and measurable outcomes.

Research focus
Applied agent evaluation
Evaluation model
Partner-specific
Program intake
By agreement
Website first published
2024-10-15
Document version
1.0.0
RESEARCH CONTROL PLANE policy · provisioning · evaluation harness · scoring control channel · provisioning, policy, telemetry PARTNER ENVIRONMENT A Dedicated domain Isolated compute Approved tools Audit telemetry access: role-based, least privilege egress: restricted routes only PARTNER ENVIRONMENT B Dedicated domain Isolated compute Approved tools Audit telemetry access: role-based, least privilege egress: restricted routes only PARTNER ENVIRONMENT C Separate access boundary configuration not disclosed access: role-based, least privilege egress: restricted routes only Environments share the control plane. They do not share data paths, identities, credentials or network routes.
Fig. 1 — Program topology. One control plane provisions and observes many engagements; the crossed links mark the connections that deliberately do not exist between participants.
§ 01 Overview

Applied Research Beyond the Laboratory

dst9.tech is an independent international research group focused on the evaluation of AI agents in controlled, task-oriented environments.

Laboratory benchmarks are useful, but they do not always reflect the operational conditions in which agents interact with files, software tools, internal workflows and human reviewers. Our research program enables selected organisations to evaluate designated AI agents in isolated environments configured for specific professional use cases.

We study reliability, task completion, tool use, error recovery, human-agent interaction and operational safety. Certain technical details of the agents, evaluation harnesses and scoring procedures remain confidential in order to protect research integrity, intellectual property and security.

Dimension 01

Reliability

Whether behaviour holds across repeated runs of the same task under the same conditions.

Dimension 02

Task completion

Whether the agreed task was finished to the acceptance criteria set in the evaluation plan.

Dimension 03

Tool use

How the agent selects, sequences and recovers from the tools made available to it.

Dimension 04

Error recovery

What the agent does after a failed step: retry, escalate, degrade or proceed incorrectly.

Dimension 05

Human-agent interaction

Where reviewers intervene, what they correct, and how the agent responds to correction.

Dimension 06

Operational safety

Whether the agent remained inside its declared boundaries throughout the engagement.

Methodological orientation

Our approach follows publicly described practice in agent evaluation: agents are run in real or isolated environments, given a limited tool set, and asked to perform verifiable tasks. Published work by AI developers and evaluation organisations, and public material on test, evaluation, verification and validation, serve as methodological reference points only. They do not indicate affiliation with, endorsement by, or cooperation with any of those organisations.

§ 02 Program

How the Program Works

Every engagement runs through the same six stages. The stages are sequential: an environment is not provisioned before its boundaries are agreed, and no evaluation is issued before the monitoring record for that run exists.

  1. Stage 1

    Environment Design

    The research team and the participating organisation define the permitted workflows, technical boundaries, evaluation period and acceptance criteria.

  2. Stage 2

    Isolated Deployment

    A separate environment is provisioned for the engagement. It may include a dedicated domain or subdomain, isolated compute resources, storage, access controls and project-specific software components.

  3. Stage 3

    Agent-Assisted Work

    Authorised users perform agreed professional tasks with the assistance of designated AI agents. Agent access is limited to the tools, files and services approved for the evaluation.

  4. Stage 4

    Monitoring and Review

    Technical events, task outcomes, intervention points and failure conditions may be recorded in accordance with the applicable agreement, security requirements and data-protection rules.

  5. Stage 5

    Evaluation

    The research team assesses whether the agent completed the intended task, remained within its operational boundaries and required human correction or escalation.

  6. Stage 6

    Environment Closure

    At the end of the engagement, access is revoked or renewed, project data is returned, retained or deleted in accordance with the applicable agreement, and the environment may be decommissioned or repurposed.

§ 03 Environments

Partner-Specific Evaluation Environments

Each engagement is logically separated from other research projects.

Depending on the evaluation design, an environment may include the components listed below. Not every component is present in every engagement; the applicable evaluation plan governs.

  • C-01 A dedicated domain or project-specific subdomain
  • C-02 Isolated virtual machines or containers
  • C-03 Controlled storage and secrets management
  • C-04 Role-based access controls
  • C-05 Restricted network routes and outbound-access policies
  • C-06 Project-specific applications and test data
  • C-07 Activity logs and evaluation telemetry
  • C-08 Human approval and emergency-stop mechanisms

Domain names and infrastructure resources may be procured and administered by the infrastructure operator on behalf of a research project. They remain technical components of the evaluation environment and do not establish ownership, corporate control or employment relationships between the operator and the participating organisation.

§ 04 Domains

Why We Use Dedicated Domains

A dedicated domain provides a clear technical boundary for a research engagement. It supports separate routing, certificates, access policies, application endpoints and audit records.

The use of a domain registered or administered by the infrastructure operator does not mean that the operator owns the participating organisation, controls its business or employs its personnel. The domain identifies an evaluation environment, not a corporate group.

A single infrastructure provider may administer environments for multiple unrelated participants while maintaining separate access rights, data boundaries and contractual relationships.

Function 01

Routing

Traffic for one engagement resolves and terminates independently of every other engagement.

Function 02

Certificates

Separate certificates and issuance records per environment, with independent renewal.

Function 03

Access policy

Identity and authorisation rules are scoped to a domain rather than to a shared estate.

Function 04

Audit records

Events are attributable to one environment, which is what makes an evaluation reviewable.

§ 05 Participation

Contractual Participation

Access to the program is provided under a bilateral research, evaluation or services agreement. It is not a free public hosting service.

The infrastructure operator procures and administers the domain, compute capacity and supporting technical services required for the evaluation environment. The participating organisation is therefore not required to contract separately with the relevant domain registrar, data-centre operator or hosting provider.

The parties' reciprocal performance may include research services, infrastructure access, participation fees, agent-assisted task execution, structured evaluation feedback and other agreed deliverables. Financial settlement is handled in accordance with the applicable agreement and may include invoicing, set-off of documented reciprocal monetary claims or another legally permitted settlement mechanism.

  • R-01 Use of designated AI agents in agreed workflows
  • R-02 Provision of task execution results
  • R-03 Preparation of structured evaluation feedback
  • R-04 Payment for research or infrastructure services
  • R-05 Supply of reciprocal services
  • R-06 Settlement by set-off of documented reciprocal monetary claims
Scope of the arrangement

A separate payment by the participant to the domain registrar or hosting provider is not required, because the operator makes those purchases. Shared infrastructure does not create a corporate relationship between participants.

§ 06 Agent use

Use of Designated AI Agents

Participation may require the organisation to use one or more designated AI agents in the workflows included in the evaluation protocol.

The agents may assist with research, software development, documentation, information processing, workflow automation, quality assurance or other approved tasks. The exact capabilities available in each engagement depend on the applicable evaluation plan.

The participating organisation remains responsible for human oversight, final decisions and compliance with its internal policies. Unless expressly agreed otherwise, an AI agent is not authorised to make legally binding decisions, communicate externally on behalf of the participant, or access systems outside the approved environment.

Terminology

Participation required the use of designated AI agents within the agreed evaluation scope. This is a contractual condition of the program, agreed in advance between the parties.

§ 07 Confidentiality

Research Confidentiality

To protect research integrity, security and intellectual property, the following are not publicly disclosed.

  • Model weights and non-public model identifiers
  • System prompts and orchestration instructions
  • Unpublished evaluation tasks and scoring logic
  • Vulnerability findings
  • Credentials, infrastructure diagrams and security configurations
  • Participant data and agent transcripts
  • Information protected by contractual confidentiality obligations
Limits of confidentiality

This confidentiality does not prevent the parties from providing competent authorities or courts with documents required by law, subject to appropriate confidentiality and procedural safeguards.

§ 08 Methodology

Evaluation Principles

Our evaluation approach is based on the following principles. Each is a condition on how an engagement is designed and run, not an aspiration.

P-01

Defined scope

Each evaluation specifies the tasks, tools, data boundaries and prohibited actions.

P-02

Environment isolation

Projects are separated through domain, identity, compute, storage and network controls appropriate to the engagement.

P-03

Human oversight

Authorised personnel review material outputs and may interrupt an agent run.

P-04

Repeatability

Where appropriate, tasks are repeated to distinguish systematic behaviour from individual-run variance.

P-05

Traceability

Relevant technical events and evaluation outcomes are recorded subject to the agreed data policy.

P-06

Least privilege

Agents and users receive only the access necessary for the approved task.

P-07

Incident handling

Unexpected behaviour, attempted boundary violations and material failures are escalated under the evaluation protocol.

P-08

Contextual assessment

Results are interpreted in relation to the environment, available tools, task design and human involvement.

§ 09 Independence

Organisational Independence

Participation in the program does not create a partnership, joint venture, agency, employment relationship, corporate group or relationship of control between the research group, the infrastructure operator and the participating organisation.

Shared technical indicators — including IP addresses, name servers, mail gateways, certificates, network routes or infrastructure domains — may result from the use of common technical services. Such indicators do not, by themselves, establish common ownership or management.

PARTICIPANT A independent legal entity PARTICIPANT B independent legal entity inferred: common ownership or control not established by shared indicators uses uses SHARED TECHNICAL SERVICES IP address ranges name servers mail gateways certificate issuers common in hosting, cloud, proxy and managed-service estates
Fig. 2 — Shared technical indicators. Both participants use the same infrastructure services; the crossed link is the inference that shared use does not support.
§ 10 Questions

Frequently Asked Questions

Who owns the project domain?

Unless otherwise stated in the applicable agreement, the domain is procured and administered by the infrastructure operator for use as part of the evaluation environment.

Does the participant pay the registrar or hosting provider?

Normally, no. The infrastructure operator procures the relevant resources and provides access under the parties' agreement. The participant's payment and reciprocal obligations are governed by that agreement.

Is access provided free of charge?

No. The program is based on contractual reciprocal performance. The exact commercial and settlement structure depends on the applicable engagement.

Does a shared IP address mean that participants belong to the same group?

No. Shared IP addresses, gateways and other network resources are common in virtual hosting, cloud, proxy and managed-service environments.

Does the research group control the participant's business?

No. The research group controls only the evaluation components identified in the agreement. The participating organisation remains legally and operationally independent.

Are details of the agents publicly available?

Not necessarily. Model configurations, system instructions, evaluation tasks, transcripts and security controls may be confidential.

Are agents allowed to operate without human oversight?

The applicable evaluation protocol defines the level of autonomy. Material business decisions remain subject to the participant's human review unless expressly stated otherwise.

§ 11 Contact

Contact

Enquiries about the research program, evaluation environments or contractual participation are handled by email. This website operates no contact form, so nothing is collected from you here beyond standard server logs.

General and program enquiries
ai@dst9.tech
Legal and data protection
ai@dst9.tech
Operator details
See the Legal Notice
Response target
Within 10 working days
Before you write

Please do not send confidential material, credentials or personal data in an initial enquiry. Where an engagement proceeds, the parties agree a confidentiality arrangement and an appropriate channel first.