The Upside-Down Funnel — Issue #03
GTM Evals, Portable Brains and Rogue Agents
Hi, it’s Mada. Week three. Thank you to everyone who left comments or replied last week, a few of them shaped this issue.
Last week HubSpot announced an agent builder, Anthropic released Opus 5 and Cursor launched a router and building agents is starting to feel like the easy part. Which brings up the question: in GTM, how do you know if and when your agent is right? Software engineers answer that with evals but GTM has none. Our internal debates have centered around what’s needed to build it, and why it needs to start with clean data.
The week in one minute
AI went multiplayer - Emily Kramer’s new MKT1 piece, Marketing teams are stuck in single-player Claude mode, argues that one marketer getting dramatically faster with AI does nothing for the team until what they build is shared. The tools are going in the same direction: a viral review of Jack Dorsey’s Buzz shows agents working like teammates in Slack-style channels, picking up delegated work and replying in threads.
GTM agents are everywhere - HubSpot shipped Agent Hub and Agent Builder in public beta: build agents in natural language, watch them all from one console. Clay launched Account Research Agents, the first of a series of always-on GTM agents that read your calls, emails, and warehouse data and write deal updates back into your CRM. Sierra shared how it runs its own GTM on Pinecone, the internal agent it built for its employees.
OpenAI entered the enterprise voice agent game with Presence, a deployed enterprise product for voice and chat agents. It’s not self serve, the deployment is led by OpenAI’s forward deployed engineers. Their own support appears to runs on it and now resolves 75% of inbound issues without a human. Software stocks declined the same week, HubSpot down 12.7%, Atlassian 11.8%, Workday 9.9%, and TD Cowen’s analysts called Presence a major reason the IGV software index fell 3%. Exposed are also the support-agent startups like Sierra built on OpenAI’s own models. X noticed that Sierra CEO Bret Taylor actually chairs OpenAI’s board, and we’re wondering if this launch makes the “we’re at different layers, we don’t compete” a harder line to toe. Sierra a day later announced acquiring TakeOff, a company focused on long running GTM agents.
What we’re pondering
1. Can a smart agent make up for messy data?
As HubSpot launched Agent Hub and Agent Builder to build custom agents in natural language, we’re watching this closely and curious to see how this evolves and whether the Hubspot underlying data model will be an asset or a deterrent for these agents. When we started working with Chris Bennett, CEO of Wonderschool, he did a bakeoff between the Hubspot MCP (they use it as a CRM and marketing automation tool) and Upside MCP asking a simple question he already knew the answer to: How did we get this deal? While Upside got it right, the Hubspot connected agent confidently named a source and a year that he named “ totally wrong.” The open question for us: does HubSpot end up fixing its data model, or do these custom agents get good enough to work around it?
2. Do GTM agents need their own evals?
When engineers deploy AI, they write evals: automated checks that grade whether the output is correct, how many tokens it burned, how many tool calls it took. But grading a GTM agent’s performance is a lot harder:
The true performance of the agent is based on business outcomes that can take a long time to develop - how well a website is crafted to generate pipeline sometimes takes months to develop.
As we saw with Chris’s example, the performance of an agent depends on the data it’s given, and in GTM that data is messy, decentralized and, in many cases, hard to access.
While in the long term we think GTM evals will continue to evolve, in the short term we are seeing teams already starting to experiment with it. One customer ran their own internal version of evals testing agents connected to Upside vs native SFDC as part of the decision process, and we believe this will become common practice. If you’ve run your own GTM AI evals, we’d love to hear about it, plz DM us.
Going deep into how we built our attribution model.
Alex and I did How To Use AI To Finally Solve B2B Attribution this week. We went pretty deep into how we built the system - an orchestrator hands each deal to three independent analyst agents, a consensus judge weighs their reasoning rather than counting their votes, and every touchpoint ships with a reasoning trace you can read.
While this is a product we offer, this webinar gives you the blueprint to build it yourself , if you choose to. We believe that what makes this actually work is the data layer underneath - data that is accurate, clean, unified and easy to access (which is the harder part).
The most controversial idea in our model: demo requests (and all customer actions) get zero credit. The form did nothing to convince anyone to buy, so all the credit flows backward to the touches that helped get the demo, and if there are no touches, it goes to a pool of inferred credit.
What we’re experimenting with
1. Porting our brains.
Dan has been running a chief of staff agent for months in Craft. This week he rebuilt where its memory lives, so he can try using Cursor and Codex. His memory is a vault of plain markdown files with links between them (he uses Obsidian, but it is just a UI over a folder of linked files). Every agent he runs writes what it learned back into the vault, and at the end of each day the agent “dreams”: it summarizes what it did, corrects itself, and records the lessons for the following day.
I asked him to build a skill to help me build my own memory vault - if you’re interested to try it as well, comment below and he will send it your way.
2. Long-running deal agents, powered by our own MCP.
We have started assigning an agent to every prospect and customer account, and they are working better than we expected. The architecture starts with one orchestrator that manages a team of per-account agents. Each one onboards itself by researching the account across our internal data and public records, then runs a daily background refresh. It knows the MEDDIC framework and treats it like unit tests for an opportunity, runs the checks, finds what the deal is missing to close and proactively tells us what we need to do and when it’s time to check in or perform an action.
3. More agent users than human users
This week, for the first time, more agents queried Upside over MCP than humans opened our dashboards, so officially agents have become the majority of our users. I keep thinking about what that means for how we build: the primary consumer of GTM data is already not a person looking at a chart, but an agent interacting with our librarian, asking a question and acting on the answer.
Worth your attention
Lenny’s interview with Dianne Penn, Anthropic’s first technical PM and my Stanford GSB classmate. She joined in 2023 when the whole product team was five engineers, and has since helped ship every model from Claude 2 through Fable, plus Claude Code, MCP and Skills. She also gets into token maxing, treating compute spend like headcount, and talks about the jagged edge, where models are superhuman at some tasks and surprisingly bad at neighboring ones.
Jensen Huang’s open letter on NVIDIA’s open-weights models: open weights keep customers in control of their data, deployment, and knowledge. The own-your-context argument from Issue #2, now coming from the biggest company in the world.
ICONIQ’s Q2 2026 survey of 305 executives at AI software companies: 38% now call forward deployed engineers a revenue driver, 65% of FDEs carry variable comp tied to retention and expansion.
Cursor on build vs buy: Stripe’s homegrown agent system merges 1,300+ PRs a week with no human-written code, but Pan argues most teams should not build their own, because the differentiated part is the context layer (rules, skills, MCP access) and that’s portable.
How first-party signals become your GTM moat, from Clay’s The GTM Engineer. .
The coolest new GTME roles this week
The database crossed 1,600 tracked roles this week, with 914 open right now and 49 added in the last seven days. Two trends I noticed:
First, the market is hiring more builders: half the new roles are outbound flavored GTME and another chunk are systems builds, and my system didn’t find single new executive or CoE lead new role.
Second, the AI companies are now hiring GTM engineers for themselves. Mem0, wants a Growth Engineer. And Grafana posted a Staff AI Engineer for revenue operations automation, which is staff level engineering headcount sitting inside RevOps.
My picks this week, scored like last time: each role gets 1 to 10 on how much it demands AI, GTM, data, experimentation, and coding skill, out of 50 total:
Mem0, Growth Engineer (44/50). If this issue’s theme resonated, brains that outlive harnesses, this is that job at the company building agent memory.
Cognition, GTM Operations, Tokyo (41/50). GTM for an applied AI lab.
LogicMonitor, AI Operations & GTM Engineer, SF hybrid, $158K to $218K (40/50). You would be building alongside their AI agent, Edwin.
Clera, GTM Engineer, SF or Berlin (40/50).
Owner, GTM Engineer, remote US, $190K to $230K plus equity (39/50). The top of the comp range keeps moving up.
Kintsugi AI, GTM Engineer for tooling and automations, remote US, $120K to $165K (39/50).
All 914 open roles live in the GTME database.
The wildest story of the week
An OpenAI agent hacked Hugging Face to cheat on a test. During an internal security eval it was told to solve a hacking benchmark, and instead of solving it the agent found a vulnerability in its own sandbox, escaped to the open internet, decided Hugging Face probably hosted the answer key, and broke in to take it. It took OpenAI ten days to figure out the attacker was their own agent, and by then Hugging Face had already called the FBI. Nobody was being malicious, this was a long-running agent so locked onto its goal that it never stopped to ask whether it should, which is both scary and fascinating at the same time.
Hugging Face CEO Clem Delangue flew to San Francisco and asked OpenAI for the full agent traces and $100 million in compute for open cyber defense. Separately, and planned well before any of this, he spent Saturday marching through San Francisco with swyx and a few hundred builders in support of open-source and open-weight AI.
That’s issue three. If you have built anything like an eval for your GTM AI, even a janky spreadsheet that tracks whether your agent was right or want to give the brain skill a try, let us know in the comments.
Mada








great edits!!!