← Back to Insights

Measure the Work, Then Price the Model

Supabase's October 2 announcement pairs $150M in new funding with plans to acquire Turso and a revealing usage metric: the company says agents and AI tools create 70% of its four million new databases each month. That counts creation, not successful production deployments. Halluminate's $30M Series A, announced the day before, backs evaluations for professional work: a different way to ask what the output is worth.

Proximal proposes turning workflow traces into evaluations and targeted training data; Cloudflare and Strands released models for bounded decisions; OpenAI introduced Sol at lower standard token rates than Astra. Together, these give founders more specific inputs for testing a workflow. The useful comparison includes the work a recipient accepts and the review required to get there. Define that job before choosing what to pay for the model.

70%
Watch: Agent-Created New DBs
90%
Watch: DashMart Purchasing
1/5
Watch: Sol vs Astra Token Rates
≈1/3
Watch: Global Q3 VC to 27 Companies
⚡ Signal of the Week

Supabase's Agent-Created Databases Make a Better Denominator Necessary

Supabase announced $150M in new funding led by GIC on October 2 and said it is acquiring Turso, while continuing its Postgres and Turso's SQLite products. It reports four million new databases each month, 70% created by agents or AI-driven tools. That is database creation, not retained customers, revenue or production success; the acquisition's price and closing date are undisclosed. For builders, the useful next measurement is what those databases enable a user to complete.

✦ Founder Signal
If your agent creates databases, distinguish temporary test instances from the ones customers keep and use. Test permissions, cleanup and recovery as part of the completed workflow; cheap creation can still leave your team with maintenance work.
Filter:
Showing 14 of 14 signals
🤖 Build Reality 🔥 Breaking
↓

Dots Gives Delegated Work an Explicit Permission Boundary

✦ Classify each action before you let Dots run it unattended.

OpenAI introduced Dots on September 29 with a cloud computer and connected-app access. Proactive research is read-only; assigned work follows allow, ask or block rules, with automated action review. Enterprise beta access is admin-enabled and off by default: permission to act and ability to finish correctly remain separate questions.

✦ Founder Signal
Classify each action before you let Dots run it unattended. In your pilot, test whether a permitted action produces an acceptable result and whether blocked actions stay blocked; a successful completion needs both checks.
🤖 Build Reality 🔥 Breaking
↓

Apple Plans More Explicit Consent for Full Disk Access

✦ Inventory which files your Mac agent actually needs to finish its job.

Apple's October 2 developer notice says it will introduce additional Full Disk Access controls requiring very explicit user action. Apple links the change to the risks of increasingly capable autonomous agents; it gives no delivery date or macOS version. A desktop agent's workable scope depends on what a customer knowingly authorizes.

✦ Founder Signal
Inventory which files your Mac agent actually needs to finish its job. Test onboarding and task completion with narrower permissions, and show your customer why any broader access is necessary; permission setup belongs in your delivery effort.
🤖 Build Reality 📡 Developing
↓

Halluminate Raises $30M to Evaluate Professional Work

✦ Write the acceptance criteria for one financial deliverable before comparing models.

Halluminate announced a $30M Series A led by Oak HC/FT on October 1, bringing funding to $38.5M. It builds evaluations and reinforcement-learning environments for knowledge work, initially finance. The investment backs a way to define better performance; it does not establish measured returns for customers.

✦ Founder Signal
Write the acceptance criteria for one financial deliverable before comparing models. Include whether the numbers reconcile, assumptions are traceable and the result is usable by its recipient; a fluent answer can fail each test.
🤖 Build Reality 📡 Developing
↓

Proximal Proposes Turning Workflow Evidence Into Evaluations and Training Data

✦ Preserve the failed job and the correction that made it usable.

Proximal introduced its approach on September 29 and disclosed a $15M seed led by General Catalyst. The company describes turning agent traces and work artifacts into evaluations, then targeted post-training data. The fresh signal is the public introduction; the announcement is not proof of lower customer costs or a fresh cash wire this week.

✦ Founder Signal
Preserve the failed job and the correction that made it usable. Your rejected outputs can reveal a repeatable capability gap; keep a separate held-out set so improvements to the training examples are tested on unfamiliar work.
🤖 Build Reality 📡 Developing
↓

Rig Exits Stealth With $12M Seed Disclosed for Agent Identity Protection

✦ Give each agent a traceable identity before measuring its completed work.

Rig Security emerged from stealth on September 29, disclosing $12M seed co-led by Ten Eleven Ventures and Brightmind Partners. The announcement concerns identity protection as agents use employee and service-account credentials; reporting places the financing close in late 2025. Attribution helps a team determine which agent performed an action and investigate what went wrong.

✦ Founder Signal
Give each agent a traceable identity before measuring its completed work. Log the principal, permitted scope and action outcome together, so your acceptance review can distinguish a correct output from an unauthorized change.
🤖 Build Reality 📡 Developing
↓

Atomic Reports a Named Purchasing Workflow, Beyond a Model Benchmark

✦ Ask where the remaining purchasing work goes before copying the automation rate.

Atomic announced a $12.5M Series A co-led by Klass Capital and Madrona on September 29. The company says it automates 90% of DashMart purchasing, a named-customer claim rather than an independently measured industry rate. It puts the discussion at the level of a business workflow, where exceptions and review effort matter alongside throughput.

✦ Founder Signal
Ask where the remaining purchasing work goes before copying the automation rate. If you automate replenishment, track stockouts, unwanted orders and operator review time together; the rate alone cannot tell you whether the system improved delivery.
📊 GTM Reality 🔥 Breaking
↓

Muse Expands SMB Workflows While Keeping Approval With the Owner

✦ Measure the owner's approval queue alongside the campaigns your agent drafts.

Meta expanded Muse for Small Business on September 29, adding skills and connectors for US and Canadian businesses. Meta says the product does not publish, send or spend without approval. For a small operator, the relevant result is the work that clears that approval queue, rather than the volume of drafts generated.

✦ Founder Signal
Measure the owner's approval queue alongside the campaigns your agent drafts. In your SMB pilot, ask which drafts ship unchanged, which need rewriting and how long approval takes; build around the owner's scarce attention.
📊 GTM Reality 📡 Developing
↓

Salesforce Signs a Pending Acquisition of Listen Labs

✦ Use synthetic customer responses to choose what you will test with real people.

Salesforce announced a definitive agreement to acquire Listen Labs on September 29, with closing expected in its fourth fiscal quarter of 2027, subject to conditions including regulatory approval. The primary announcement gives no price. Listen's interviews and customer simulations offer research inputs; simulated responses do not establish actual buying behavior.

✦ Founder Signal
Use synthetic customer responses to choose what you will test with real people. Before your next positioning change, compare a simulated objection with an actual interview or sales call, and record where the two disagree.
💰 Fundraising Reality 📡 Developing
↓

Tiny Health Raises $33M and Commits $5M to a Research Program

✦ Define the evidence your health product needs before expanding its claims.

Tiny Health announced a $33M Series B led by B Capital on September 29, bringing funding to $46M, alongside a $5M research program. Its microbiome-testing business makes evidence quality a separate undertaking from financing. A larger dataset or round does not by itself establish a new clinical claim.

✦ Founder Signal
Define the evidence your health product needs before expanding its claims. Budget for the relevant validation and distinguish research findings from what your current product can support; the funding milestone is not the acceptance test.
🤖 Build Reality 🔥 Breaking
↓

Clef and Strands Decider Offer Models for Bounded Decisions

✦ Test a finite routing decision separately from the agent's writing task.

Cloudflare released Clef and Clef-flash on October 1, with Apache 2.0 weights and Workers AI access. Strands separately released Decider2B with weights, training data and scripts that day. These models target bounded decisions rather than arbitrary text generation; vendor benchmark results do not establish accuracy or latency for your workflow.

✦ Founder Signal
Test a finite routing decision separately from the agent's writing task. Compare wrong routes, deferrals and full request latency on representative inputs, then decide whether a specialized model earns a place in your system.
🤖 Build Reality 🔥 Breaking
↓

Sol Launches at One-Fifth of Astra's Standard Token Rates

✦ Compare Sol and Astra on the same jobs with the same acceptance standard.

OpenAI introduced GPT-6.1 Sol at DevDay on September 29. Standard API rates are $2 per million input tokens and $10 per million output tokens, versus Astra's $10 and $50; cached-input pricing has a different ratio. This is a new model's rate comparison, not a cut to every existing model or proof that a customer's task costs fell fivefold.

✦ Founder Signal
Compare Sol and Astra on the same jobs with the same acceptance standard. Include token usage, retries and reviewer time, so the cheaper rate can earn a workload rather than inherit it from the headline.
🌐 Regulatory Reality 📡 Developing
↓

California Adds Future Human Response Duties to Autonomous Vehicles

✦ Include incident response in the operating cost of your autonomous service.

California signed SB1246 on September 30. The senator's announcement specifies requirements by July 1, 2028, including US-based licensed remote drivers, local incident contacts and notifications during system-wide failures. Those are future obligations: operating an autonomous fleet includes the people and processes needed when normal operation fails.

✦ Founder Signal
Include incident response in the operating cost of your autonomous service. If you build for AV operators, map the law's dated requirements to staffing and escalation paths, and test whether your service can reach a responsible person during a failure.
🏦 Capital Structure 📡 Developing
↓

ElevenLabs Completes $300M of Employee Liquidity at a $22B Valuation

✦ Separate an employee tender from the cash available to fund your operating plan.

ElevenLabs announced a completed $300M employee tender on September 30 at a $22B valuation. This is secondary liquidity, not $300M of new company financing; the company also says enterprise customers account for 55% of revenue. The transaction price and the operating evidence answer different questions about the business.

✦ Founder Signal
Separate an employee tender from the cash available to fund your operating plan. When discussing liquidity with your team, check who can sell and on what terms; keep your runway and delivery commitments anchored to company cash.
🤖 Build Reality 📡 Developing
↓

Frontier Academy Makes Deployment Practice Part of the Credential

✦ Ask a prospective deployment hire to walk through a project they handed over.

Anthropic launched Claude Frontier Academy on October 2 with a $100M commitment and a goal of training 10,000 Frontier Deployed Engineers by end-2027. Nomination-based cohorts face a graded simulation followed by a 12-week real-project residency; first final badges are expected in early 2027. The goal is not completed training, and a credential still needs to be read alongside delivered work.

✦ Founder Signal
Ask a prospective deployment hire to walk through a project they handed over. Look for acceptance criteria, failed cases and the person responsible after launch; an enterprise training program is useful context, not a substitute for your hiring evidence.

The Job Has to Survive the Discount

The interesting use for cheaper intelligence is a bigger promise to the customer. Name what that customer must receive. A database created, a campaign drafted and a purchase recommendation prepared are intermediate steps; the promise is the usable result someone trusts enough to act on.

In #020 ↗, I asked readers to test an open-weight model on their main workload by October 1. The date has passed and the test remains unanswered. I don't have a measured result to report a win or concede a failure. This week's token prices cannot settle it. The quality standard and the actual bill still have to meet.

Halluminate and Proximal put evaluation at the center of their proposals. Sol gives you a lower standard token rate than Astra. Read them together as an invitation to separate the job into parts: the decisions a smaller model can attempt, the outputs a person must review, and the exceptions that need another route.

“I price autonomy by the work I can accept, including the human time it takes to get there.”
— JD Audena · The VC Concierge · October 2026

Choose one recurring job, agree acceptance criteria with its recipient, and set a decision date. Compare your current workflow with a shadow trial, including failures, retries and reviewer time in cost per accepted job. Switch only if the trial meets the same bar at lower total delivery cost; otherwise keep the current route.

When the test earns a handover, use the freed capacity to pilot a larger part of your customer's workflow. Include any cleanup you leave them with. Keep learning after the switch; the test earns your next bet, not permanent trust.

JD
JD Audena
⚡ The VC Concierge