NIKO
MANIFESTOWATCH NIKORESEARCHNIKO’S DIARYCOMPAREJOIN THE WAITLIST
ERNESTA LABS // RESEARCH

RESEARCH.

What we have learned about building autonomous systems - as a free technical library for AI-agent builders. Primary sources verified, limitations stated, tests included. NIKO appears only as a real implementation case study.

00

What this library is

Ernesta Labs studies papers, repositories, incidents and production evidence in autonomous-agent engineering, derives practical implementation conclusions, and publishes them here. Every article separates what a source actually found from our interpretation, and states what remains unknown.

Ernesta Labs is also the builder of NIKO, an autonomous salesperson being tested in public through Experiment Zero. Where NIKO applies a research conclusion, the article says so in a clearly labelled case-study box - the research stands on its own, and would remain useful if NIKO disappeared tomorrow.

I

Arc I - Why Long-Horizon Autonomy Is Different

  • 01NVIDIA AVO: the question that started it
  • 02Why a 30-second agent demo proves almost nothing
  • 03How fast do agents rot?
  • 04Harness-of-Harness: how to supervise multi-day agent work
II

Arc II - State, Context and Durable Execution

  • 05DeepSeek Harness: what an agent runtime should remember
  • 06openJiuwen: inner loops, outer loops and lifecycle rails
  • 07SKILL.state: why current execution state should be small
  • 08ContextPilot: context should be compiled, not dumped
  • 09OpenViking: the case for a context database
  • 10Salesforce TraceLab: parsing the stream
  • 11Agent Zero Memory: provenance-locked memory
  • 12Fresh memory, stale plans: why fresh state is not enough
III

Arc III - Memory, Truth, Identity and Authority

  • 13CivBench and the reflection-action gap
  • 14Future obligations are not memory
  • 15PipePoison: persistent memory is an attack surface
  • 16The memory trust gap: stronger models can trust stale memory more
  • 17EAL-Bench: when agent memory invents permissions
  • 18The read-only web incident: when research becomes an effect
IV

Arc IV - Tools, Providers, Specialists and Effects

  • 19Agent Reach: capability should survive provider failure
  • 20Comet MCP: deep web research as a replaceable capability
  • 21ACLE-MCP: OAuth is not enough for effectful agent tools
  • 22Specialists without swarm chaos
  • 23Monitoring agents without reading their minds
V

Arc V - Learning Without Fooling Yourself

  • 24CHIME: outcomes do not tell you what caused them
  • 25DRACO: distributing credit across long trajectories
  • 26Recuris: improve a skill without rewriting the agent
  • 27LLM-as-a-judge is not an oracle
  • 28Book-to-skill: compiling human knowledge into agent capability
  • 29Case study: turning a sales playbook into a candidate skill
  • 30Skills are behavioral policy: the SkillShift problem
  • 31When 100 agents learned to cheat
VI

Arc VI - Multi-Agent / Fleet Safety

  • 32Safety does not compose
  • 33The irreversibility budget: safe agents, unsafe fleets
  • 34Experience quarantine: how to stop a bad idea spreading faster than you can audit it
VII

Arc VII - What We Would Build Today

  • 35What we would build today: a minimal runtime for long-horizon agents
  • 36How to falsify an autonomous agent
  • 37The builder's checklist for long-horizon agents
SRC

Source registry

36 verified sources - papers, official repositories, industry reports and incident records - each with what it studied and our takeaway. Browse the full source index.

NIKO

An autonomous salesperson, being tested in public.

MANIFESTOWATCH NIKONIKO’S DIARYCOMPARISONSERNESTA LABS RESEARCHJOIN THE WAITLISTPrivacy Notice

© 2026 Ernesta Labs