Kinesthetic

Agents that learn by doing

AI is not just hype. In recent years, we have made massive progress in building models that contain vast knowledge and can effectively perform complex tasks, and we expect this will continue.

While this will yield useful general intelligence, many tasks and domains require agents to possess a deep understanding of large and complex knowledge that can never exist in foundation training data to be useful. However, extending agents beyond foundation capability today presents challenges that prevent AI from reaching its full potential in these specialized domains.

Our work revolves around how agents can learn by doing, and how humans can shape agent behavior through processes that feel more like continuous improvement through collaboration than offline data generation efforts. Ultimately, we seek the optimal method for humans to endow agents with new learnings that remain usable, auditable, and portable.

What are we building?

We enable continuous agent improvement by leveraging optimized processes around data already available for capture and human-agent collaboration to enable agents to become experts in specialized domains.

We have research expertise on having agents learn from any feedback they receive from users, engineers, operators, and their own experiences. And by implementing ergonomic interfaces on top of these methods, we enable agents to improve at a rate significantly faster than today's go-to processes such as making ad-hoc system changes or building out extensive RL environments.

A natural byproduct of this is owning portable, auditable knowledge that can be transferred across any model or agent in your ecosystem.

Who is this for?

Companies building AI agents (harness for non-deterministic workflows) for complex, out-of-distribution tasks that require large amounts of domain specific knowledge and standard procedures that are unmanageable via manual curation.

The team operates the agent in production and captures a volume of feedback that outpaces their ability to fully utilize it.

Leadership is looking for improvements on the agent, and more specifically wants to see the rate of improvement/capability extension increase.

How is our work different from other continual learning neolabs?
  • Beyond verifiability: we are primarily interested in combating failure modes beyond those that can be optimized against verifiable rewards. The limitations preventing agents from delivering on promises of transformative change in specialized domains typically can't be naturally represented with verifiable reward signal.
  • Token space: we are researching methods that keep knowledge in natural language representation, but can optionally be distilled into weights. We believe that auditability and portability are key operational requirements for trustworthy agents and investable efforts.
  • Accessible: we focus on utilizing data from unstructured feedback and the agent's historical experience. For humans to be able to teach agents to do complex work in specialized domains, we need to enable end-users, operators, and experts to shape agent behavior without extra steps. Removing the barriers to entry of ML expertise and up-front cost is necessary to enable the long tail of specialized domains to unlock intelligence that can understand how they work.
How can we get in touch?

Send us an email!