Reliability
Whether behaviour holds across repeated runs of the same task under the same conditions.
Independent international research group
We operate partner-specific evaluation environments where AI agents can be tested in controlled professional workflows, with defined access boundaries, human oversight and measurable outcomes.
dst9.tech is an independent international research group focused on the evaluation of AI agents in controlled, task-oriented environments.
Laboratory benchmarks are useful, but they do not always reflect the operational conditions in which agents interact with files, software tools, internal workflows and human reviewers. Our research program enables selected organisations to evaluate designated AI agents in isolated environments configured for specific professional use cases.
We study reliability, task completion, tool use, error recovery, human-agent interaction and operational safety. Certain technical details of the agents, evaluation harnesses and scoring procedures remain confidential in order to protect research integrity, intellectual property and security.
Whether behaviour holds across repeated runs of the same task under the same conditions.
Whether the agreed task was finished to the acceptance criteria set in the evaluation plan.
How the agent selects, sequences and recovers from the tools made available to it.
What the agent does after a failed step: retry, escalate, degrade or proceed incorrectly.
Where reviewers intervene, what they correct, and how the agent responds to correction.
Whether the agent remained inside its declared boundaries throughout the engagement.
Our approach follows publicly described practice in agent evaluation: agents are run in real or isolated environments, given a limited tool set, and asked to perform verifiable tasks. Published work by AI developers and evaluation organisations, and public material on test, evaluation, verification and validation, serve as methodological reference points only. They do not indicate affiliation with, endorsement by, or cooperation with any of those organisations.
Every engagement runs through the same six stages. The stages are sequential: an environment is not provisioned before its boundaries are agreed, and no evaluation is issued before the monitoring record for that run exists.
The research team and the participating organisation define the permitted workflows, technical boundaries, evaluation period and acceptance criteria.
A separate environment is provisioned for the engagement. It may include a dedicated domain or subdomain, isolated compute resources, storage, access controls and project-specific software components.
Authorised users perform agreed professional tasks with the assistance of designated AI agents. Agent access is limited to the tools, files and services approved for the evaluation.
Technical events, task outcomes, intervention points and failure conditions may be recorded in accordance with the applicable agreement, security requirements and data-protection rules.
The research team assesses whether the agent completed the intended task, remained within its operational boundaries and required human correction or escalation.
At the end of the engagement, access is revoked or renewed, project data is returned, retained or deleted in accordance with the applicable agreement, and the environment may be decommissioned or repurposed.
Each engagement is logically separated from other research projects.
Depending on the evaluation design, an environment may include the components listed below. Not every component is present in every engagement; the applicable evaluation plan governs.
Domain names and infrastructure resources may be procured and administered by the infrastructure operator on behalf of a research project. They remain technical components of the evaluation environment and do not establish ownership, corporate control or employment relationships between the operator and the participating organisation.
A dedicated domain provides a clear technical boundary for a research engagement. It supports separate routing, certificates, access policies, application endpoints and audit records.
The use of a domain registered or administered by the infrastructure operator does not mean that the operator owns the participating organisation, controls its business or employs its personnel. The domain identifies an evaluation environment, not a corporate group.
A single infrastructure provider may administer environments for multiple unrelated participants while maintaining separate access rights, data boundaries and contractual relationships.
Traffic for one engagement resolves and terminates independently of every other engagement.
Separate certificates and issuance records per environment, with independent renewal.
Identity and authorisation rules are scoped to a domain rather than to a shared estate.
Events are attributable to one environment, which is what makes an evaluation reviewable.
Access to the program is provided under a bilateral research, evaluation or services agreement. It is not a free public hosting service.
The infrastructure operator procures and administers the domain, compute capacity and supporting technical services required for the evaluation environment. The participating organisation is therefore not required to contract separately with the relevant domain registrar, data-centre operator or hosting provider.
The parties' reciprocal performance may include research services, infrastructure access, participation fees, agent-assisted task execution, structured evaluation feedback and other agreed deliverables. Financial settlement is handled in accordance with the applicable agreement and may include invoicing, set-off of documented reciprocal monetary claims or another legally permitted settlement mechanism.
A separate payment by the participant to the domain registrar or hosting provider is not required, because the operator makes those purchases. Shared infrastructure does not create a corporate relationship between participants.
Participation may require the organisation to use one or more designated AI agents in the workflows included in the evaluation protocol.
The agents may assist with research, software development, documentation, information processing, workflow automation, quality assurance or other approved tasks. The exact capabilities available in each engagement depend on the applicable evaluation plan.
The participating organisation remains responsible for human oversight, final decisions and compliance with its internal policies. Unless expressly agreed otherwise, an AI agent is not authorised to make legally binding decisions, communicate externally on behalf of the participant, or access systems outside the approved environment.
Participation required the use of designated AI agents within the agreed evaluation scope. This is a contractual condition of the program, agreed in advance between the parties.
To protect research integrity, security and intellectual property, the following are not publicly disclosed.
This confidentiality does not prevent the parties from providing competent authorities or courts with documents required by law, subject to appropriate confidentiality and procedural safeguards.
Our evaluation approach is based on the following principles. Each is a condition on how an engagement is designed and run, not an aspiration.
Each evaluation specifies the tasks, tools, data boundaries and prohibited actions.
Projects are separated through domain, identity, compute, storage and network controls appropriate to the engagement.
Authorised personnel review material outputs and may interrupt an agent run.
Where appropriate, tasks are repeated to distinguish systematic behaviour from individual-run variance.
Relevant technical events and evaluation outcomes are recorded subject to the agreed data policy.
Agents and users receive only the access necessary for the approved task.
Unexpected behaviour, attempted boundary violations and material failures are escalated under the evaluation protocol.
Results are interpreted in relation to the environment, available tools, task design and human involvement.
Participation in the program does not create a partnership, joint venture, agency, employment relationship, corporate group or relationship of control between the research group, the infrastructure operator and the participating organisation.
Shared technical indicators — including IP addresses, name servers, mail gateways, certificates, network routes or infrastructure domains — may result from the use of common technical services. Such indicators do not, by themselves, establish common ownership or management.
Unless otherwise stated in the applicable agreement, the domain is procured and administered by the infrastructure operator for use as part of the evaluation environment.
Normally, no. The infrastructure operator procures the relevant resources and provides access under the parties' agreement. The participant's payment and reciprocal obligations are governed by that agreement.
No. The program is based on contractual reciprocal performance. The exact commercial and settlement structure depends on the applicable engagement.
No. Shared IP addresses, gateways and other network resources are common in virtual hosting, cloud, proxy and managed-service environments.
No. The research group controls only the evaluation components identified in the agreement. The participating organisation remains legally and operationally independent.
Not necessarily. Model configurations, system instructions, evaluation tasks, transcripts and security controls may be confidential.
The applicable evaluation protocol defines the level of autonomy. Material business decisions remain subject to the participant's human review unless expressly stated otherwise.
Enquiries about the research program, evaluation environments or contractual participation are handled by email. This website operates no contact form, so nothing is collected from you here beyond standard server logs.
Please do not send confidential material, credentials or personal data in an initial enquiry. Where an engagement proceeds, the parties agree a confidentiality arrangement and an appropriate channel first.