At Discovered Labs we work with 10M to 50M USD ARR companies to help them get more leads, users and customers from Google, Bing and AI assistants such as ChatGPT, Claude and Perplexity. We approach marketing the way engineers approach systems: data in, insights out, feedback loops everywhere. We are a deeply technical team; you will work directly with founders who can go deep on architecture, code, and product. You'll join the Automations team at Discovered Labs. We build the agent swarm: a growing fleet of agents and workflows that do real client work end to end, and the engineering that makes them trustworthy. Your mandate is to expand what the swarm can do, and prove that what it does is right: evals that fail a build, traces that explain a decision, guardrails that hold when a model doesn't. What You'll Do: - Agent eval harnesses: golden suites, deterministic checks, and online scoring that gate every agent change in CI. - Agent observability: step traces, token and cost accounting, run-level SLOs, and failure taxonomies. - Data lineage and provenance: every claim an agent makes should be traceable to its source. - End-to-end agents with human review, with a curated review surface so an expert can approve, correct, or reject. - Interfaces for non-engineers: dashboards and controls that let the SEO team run, inspect and trust agent work. - Expanding what agents can do: pull requests against real repositories, OAuth authentication, operating real sites and CMSs through the browser. - Closing the loop: turn AI perception and citation data into strategy recommendations and measurable outcomes. - Algorithms and scoring models: internal linking and semantic relevance scoring, thread scoring, content and citation quality. - Shared, composable modules and scaling workflows across many clients.
- 5+ years in software engineering, with meaningful recent time on LLM-backed or ML-backed production systems, including 2+ owning a production system end to end. - Python, React, TypeScript and strong systems fundamentals. - LLM application engineering in production: agents, tool use, structured outputs, retrieval, prompt architecture. - Evaluation systems for non-deterministic systems: golden datasets, regression gates, offline and online scoring. - Observability for AI systems: tracing, run inspection, cost and token accounting (Langfuse, LangSmith, Braintrust, OpenTelemetry or equivalent). - Debugging non-determinism. - Pipeline orchestration: Airflow, Dagster, Temporal or similar. - Third-party API integration: auth flows, rate limits, pagination. - Own your infrastructure: containers, CI/CD, deployment, monitoring, credential management. - Product sense for expert users and strong written communication. Preferred Qualifications: - Browser automation or authenticated web agents (Playwright, Puppeteer, computer-use models) - Applied ML: ranking, scoring, classification, or recommendation systems in production - Terraform, Kubernetes - Prior experience at a fast-moving startup