Storrow Labs · Labs · Updated July 2026

Emerging tools, evaluated before they become defaults.

We prototype new agent frameworks, models, and deployment patterns against real operating requirements: data access, permission boundaries, tool execution, evaluation, observability, human review, and recovery.

OpenClaw · Hermes Agent · Thinking Machines · Open models · Sandboxed agents · Evaluation

Current research

Frontier capability with an explicit production boundary.

Systems under evaluation include OpenClaw, Hermes Agent, and models and tooling from Thinking Machines. A Labs prototype is not automatically a production recommendation. Technologies move into client environments only when their operating and security characteristics fit the approved use case.

01

Agent boundaries

Who may invoke the agent, which tools it may call, what it can read or change, and where human approval is required.

02

Deployment isolation

Identity, secrets, network access, sandboxing, tenant boundaries, logging, rollback, and the difference between a prototype and a production environment.

03

Model & workflow evaluation

Test representative work for quality, failure modes, prompt injection, data exposure, operating cost, latency, and recoverability.

Most AI advice begins with the model. We begin with the boundary: who can invoke it, what data can reach it, what it can change, and how a human can stop it.

Have a frontier workflow worth testing?

We will define the operating boundary before selecting the technology.

Discuss a pilot →