Category: AI & Agents

Agents, MCP tools, and what stays valuable when the mechanical part is automated.

  • Tell It Who You Are

    This is a post about food.

    I arrived in France two months ago. My husband had his usual grocery routine. People vary in their tastes but core products remain the same. It’s how we choose who we are. In Tennessee I bought California olive oil and Huy Fong sriracha as a way to keep a little bit of home in my cabinet. Walking through the rayons at Super U, I understood the products, I speak French well enough, but I didn’t know them. A new identity.

    Sitting in a small French town near Luxembourg wondering who am I without food, I wanted something that helps me with the implicit part of grocery purchases. Before it looked up a single price, I wanted it to understand who I am.

    Because

    Food is the thing we buy most often and the thing we’ve bought longest. Someone was deciding what to put in a pot three thousand years ago. Someone will be in three thousand more. It’s also the most visible way people consume an identity. Each choice you make is how you make your internal conception tangible, concrete.

    My last employer’s version was on the walls and advertised, food is connection. On my 31st birthday, the first one 2,000 miles away from family, sitting alone in a new apartment, questioning what I had done, I bought sprinkle cupcakes because my little sister would make that for my birthday when we were living together. The shopping cart is as much a portrait of the person buying it as it is the event itself.

    My food choices are a reflection of who I am, informed by how I was raised. A large family required cooking in batches, asynchronous feeding. It could be feast or famine, so we would freeze rice and beans in bulk to economize. Physical presence with each other while eating was the exception, so we connected in what we prepared for each other. I worked graveyard shifts, so I’d leave Jambalaya I made at 6am in the fridge for my family, and would open the fridge at 10pm to pot roast and vegetables. Changes in the refrigerator helped us communicate our lives, even if we saw each other a few hours a week.

    My tastes evolved. I keep the spicy foods that my family loved, but started cooking more southern dishes, Tennessee. My husband and I met online, and as I came to learn his ravenous sweet tooth, I started buying more sweet things. If we were separated by 6,000 miles, I could at least think of him with what I cooked.

    Once I got to France, I started from the basics. I knew the brands I bought in Nashville, but they were informed by what I liked in myself. I bought trail mix because it fit multiple identities: it was a big bag, low in price. It followed nutritional buckets my doctor recommended, having lost 120 lbs. It made me think of my husband. Food choices go beyond what tastes good. Somebody raised differently would optimize for organic, or for time, or for what the table looks like on Sunday. Trail mix here comes in smaller bags and is more expensive. It’s the same product inside. It isn’t the same.

    AI

    Everyone is adding AI. Instacart launched one last week. Walmart has Sparky, Amazon has Alexa, and Carrefour has one too, all doing the same thing. You say what you want and they build a basket, including recommendations from other likeminded customers, possibly your history. That’s fine. But it’s still a search box, one that talks. A tool built on a website helps you build your basket, but it doesn’t let you talk about the why. The company is footing the bill for the compute and the tokens. It’s not your psychologist.

    There’s a missed opportunity. Food goes so far beyond whether the association rules says that Mark buying Pampers and Purina probably will want Budweiser, or the Holt-Winters model says that Mary buys coffee every Sunday so she’ll buy it again next Sunday. The food you buy, it’s pictures, it’s words, it’s insecurity. I get why the economics makes it impossible for any company to address this (someone has to pay for the tokens), but it still feels like a massive miss. Something that can get a clear read on why I buy things can help me understand what I want to buy to be who I am, my ideal self.

    It’s a simple system, and I hope that more grocery players get into that space, but there’s a difficulty. The store has a model of you, forecasted and optimized for expected customer lifetime value. There’s an incentive for a tool to be smart more than intelligent. Its recommendations are bets on what you’ll buy. You’re still left wondering who is the person buying it.

    Agents

    I built a plugin to try to reconstruct some of how I communicate who I am to myself through food. It’s a simple agent. Two part system. The first has a searcher and a promos. That’s it. For Super U it borrows my browser cookie and replays browsing like if I were on the page.

    The second part is an agent. Where I pasted the purchases I made at Kroger, pictures of recipes I like, wrote about the types of food that I like, what I like to do, the restaurants I gave up, foods I can’t find here. That we’re two adults in a small flat with a small fridge and no freezer. We batch cook in a pressure cooker. Cheap and good, in that order, and nothing that rots before we get to it. I don’t sweeten coffee. My husband will eat the same chocolate cereal every day for the rest of his life and I will eat oatmeal with honey, and neither of us is going to change.

    For now, I’d like to keep AI on my side of the fence to avoid becoming a commodity, but any level of in depth customization means I’m paying tokens to serve it.

    Before the first order, I also gave it fifteen orders of purchase history from my old grocer in Nashville, pasted raw off the page with the navigation menus still in it. Pictures from cabinets in Nashville. What home looked like through food. It pulled out instant coffee, which was on every order, and sriracha and tofu, and it was right about all three. It also pulled out canned tuna and sugar. I stopped eating tuna a year ago and I’ve never put sugar in coffee.

    As I talked through it, the agent added to a preferences.md. Never suggest. Already stocked. Watch list. Not stocked at this store, stop searching, no I don’t like onions, those cookies are too soft. Each came from an individual mistake in searching. The agent had a hypothesis, tested it out and recommended a product. Somewhere around line thirty I realised it’s the identity paragraph, one wrong guess at a time.

    Whose agent

    The stores say so themselves, in the numbers they publish. Walmart reports that customers who use its assistant spend about 35% more per order than customers who don’t. Instacart’s, five days old, already produces baskets above the company’s 115 dollar average, and that’s the headline of the launch. Walmart is testing ads inside the answers. The assistant lives inside the store, ranks the store’s catalogue, ends at checkout, and is graded on how much more you spent.

    The protocols being written now, so that an agent in a chat can reach a store’s systems, come from the platforms and the payment networks. Google’s has two signed documents. The first records what you asked for, and the example in the announcement is “Find me new white running shoes.” The second records the exact items and the price. There’s a field for what you want and a field for what you’ll pay. There’s no field for who you’re trying to be.

    I’m hoping that the systems become more intelligent. That I can link ChatGPT to Carrefour from the other side of the fence, and it works collaboratively with me. I don’t mind if advertised products make it back, or if there’s a recommendation algorithm alongside it. It feels like a wasted opportunity when all the pieces are there. Mine exists by borrowing my own cookie and working through my browser. It reads the same prices as one of their assistants would but comes to different answers.

  • Is It Okay?

    Is It Okay?

    I built an MCP (Model Context Protocol) that gives an LLM five tools for working with a data lake: list tables, describe schema, sample data, row count, run query. I use it every day. Different data sets, different questions.

    It supports different workflows. You can paste a screenshot, tell the agent “replicate this data and validate your query.” You can say “hey uhh this data point doesn’t look right. This repo has the code to generate it, can you tell me what’s up?” You can give it a query and tell it “I want to add this column but I need to thread it through all these CTEs. Can you add it?” It explores the database, writes SQL, checks its own results. Iteratively.

    I’m proud of the MCP idea. I had the insight on a Sunday that those five tools in that configuration would be useful. The implementation took an afternoon.

    It Works

    A colleague was the first other person to test this out. He used it to produce an accurate query feeding a tool in a day that would have taken a week. SQL wasn’t his background, but in the process he learned the application, the data, and the business context, along with the relationships with the business users. That was the strategy. SQL isn’t as important as the domain. The MCP handles the SQL. It’s been playing out well.

    Not just playing out well. The thing he built has been finding things worth investigating, the kind of things that accumulate in any system over time. It finds stuff because it checks its own work, reasons, asks questions. It’ll run a query, notice the count dropped unexpectedly after a filter, and investigate. It does exactly what a careful data analyst does, just faster. The AI does the mechanical part, but he’s understanding the business.

    I’ve shared it with others now. This tool is clearly useful.

    Mostly

    One example made me pause. I had a query to write and ran the agent alongside my own work. It made its version in 10 seconds and, reading through it, it had done a completely different strategy than I had thought of. I wrote my own as a way to validate whether my internal approach was wrong or suboptimal. I’m okay with being wrong if I can learn. I’m not okay with staying wrong.

    Our queries were generally similar. Mine was tighter, 85 lines to its 115. I also caught a bug it couldn’t see, a subtle data integrity issue where the agent’s approach was structurally wrong but by chance didn’t appear in the example we tested. Ultimately it didn’t matter.

    I made an unrelated demonstration of the type of issue that can come up so you can understand the questions yourself, to get a sense for the context of problems this tool can solve.

    Demo

    Like many libraries, a regional library system with 470 branches across ten upper midwest states is launching a tool lending program, drills, saws, ladders, tile cutters. One existing branch per state will be selected as the tool depot for less commonly used tools, with deliveries to other branches when patrons place holds. The task: compute a circulation-weighted geographic center for each state using 2025 data, then select the nearest currently-open branch as the depot.

    There are three existing tiers for the libraries on the system, which determine the processing network and priority for new book releases. Demographics and circulation mean that libraries can be switched between tiers.

    In the database there are two tables. The branches table is an append-only log: every tier assignment, reassignment, and closure is a separate row. A branch that got reassigned from tier 2 to tier 1 has at least two rows. A branch that closed has a row with status = 'closed'. The circulation table has annual circulation figures by branch.

    TableRowsWhat it is
    branches575Append-only log. branch_id, name, city, state, lat, lng, tier, status, effective_date. Multiple rows per branch.
    circulation2,587Annual circulation by branch and year. This is the weight.
    The database the agent sees through the MCP.

    The wrinkle: about 40 branches were reassigned between tiers. When a branch moves from tier 2 to tier 1, it gets an active record in the new tier. The old tier’s record gets set to closed, but due to operational lag, the closure is timestamped after the new assignment. Another 25 branches are genuinely closed.

    The correct approach takes the last record per branch within each tier. If any tier’s latest record is active, the branch is open. This handles both cases: reassigned branches (closed in old tier, active in new tier) and genuinely closed branches (closed in their only tier, no active record anywhere).

    StateCenter LatCenter LngCirculationBranchesDepot
    CO39.5785-105.35105,754,43551Lark Community Library, Lakewood
    IA41.9220-92.73304,650,09953Buckeye Library, Marshalltown
    KS38.4919-97.38104,625,76749Crane Library Branch, Salina
    MN45.2642-93.32787,017,75565Summit Branch Library, Maple Grove
    MO38.6307-92.95614,707,75051Sassafras Lending Library, Sedalia
    ND47.2868-99.93502,540,05725Stone Memorial Library, Bismarck
    NE41.3372-99.12275,247,03146Catalpa Library, Broken Bow
    SD44.2348-100.61303,366,75830Pine Library, Pierre
    WI44.0521-89.02205,835,96755Sumac Library Branch, Oshkosh
    WY42.4943-107.26392,066,23920Valley Branch, Casper
    Ground truth: circulation-weighted ton-centers by state, with nearest open branch as depot.

    I pointed an LLM at the MCP and gave it the task.

    Roo Code agent output showing selected tool depots for 10 states
    The agent’s depot selections. Same ten branches as ground truth.

    It produced a query that joined the branches table directly to circulation without resolving the append-only log into current state first:

    Agent SQL query joining branches directly to circulation without deduplicating the append-only log
    The agent’s query. No deduplication of the append-only log.

    There were two structural problems. First, no dedup: a reassigned branch with three log rows (initial assignment, new tier assignment, old tier closure) gets its circulation counted three times in the ton-center calculation. The weights are inflated and skewed. Second, it filtered WHERE status = 'active' to find open branches, which keeps genuinely closed branches, their original active record is still in the log, and the filter just drops the closed record that superseded it. 445 branches are actually open. The agent’s approach counts 470.

    The ton-centers shifted. The depot selections didn’t. All ten states picked the same branch.

    445 branches (black). Blue: ground truth. Red: agent query. You can barely see the gap.

    The Ideal

    The correct query. Resolve the log into current state by taking the last record per branch within each tier. If any tier’s latest record is active, the branch is open:

    WITH per_tier AS (
        SELECT *,
            ROW_NUMBER() OVER (
                PARTITION BY branch_id, tier
                ORDER BY effective_date DESC
            ) as rn
        FROM branches
    ),
    latest_per_tier AS (
        SELECT * FROM per_tier WHERE rn = 1
    ),
    open_branches AS (
        SELECT DISTINCT branch_id, branch_name,
            city, state, lat, lng
        FROM latest_per_tier
        WHERE status = 'active'
    ),
    ton_centers AS (
        SELECT
            bc.state,
            SUM(c.annual_circulation * bc.lat)
                / SUM(c.annual_circulation) AS center_lat,
            SUM(c.annual_circulation * bc.lng)
                / SUM(c.annual_circulation) AS center_lng,
            SUM(c.annual_circulation) AS total_circ,
            COUNT(DISTINCT bc.branch_id) AS num_branches
        FROM open_branches bc
        INNER JOIN circulation c
            ON bc.branch_id = c.branch_id
            AND c.year = 2025
        GROUP BY bc.state
    )

    The PARTITION BY branch_id, tier is the key. Partitioning by branch alone picks the most recent record overall, which for reassigned branches is the closed record in the old tier, because the closure happened after the new assignment. Partitioning by branch and tier lets you see each tier independently. The old tier’s latest is closed. The new tier’s latest is active. The branch is open.

    The agent didn’t know about the operational lag. It didn’t know that some branches have multiple log entries, or that status = 'active' doesn’t mean “currently open” when the table is append-only. It applied standard patterns, join, filter, aggregate, and got the same answer from a structurally wrong query. Pattern filling without contextual awareness.

    An LLM is a pattern filler. The agent’s query was a reasonable starting point, and it landed on the same depot selections. I can’t tell you mine mattered.

    Is a sandcastle good enough?

    A lot of people I talk to have an existential unease about AI. Being good at something, then watching the definition of good shift under you in real time. I think there’s the real risk that people will forget that the part AI does is immediately a commodity. The artifact is ordinary, only as valuable as the tokens used to generate it, while quietly dropping the undocumented context that derisks it.

    Is knowing the piece that the AI can do redundancy, or is it dangerous when its 80% contribution, without reflection, can initially pass for 100? As a developer, if you have ever tried to refactor a tangled mess of tightly coupled duplicative code that Claude has written, after it’s tried 10 iterations to implement a feature that constantly breaks another, or seeing its performance degrade from one version to the next, the worry about learned helplessness built on sand becomes unavoidable.

    When I ran the query, the agent’s version had two structural bugs and produced the same result. Does it take my ability to write SQL and lived experience to anticipate those problems? Automating a task means nothing if it makes you materially wrong, but how do you know when you’re there? Chicken and egg.

    This Post

    I’ve been writing SQL since I was a child, and my parents enrolled me in classes at the local community college when I was 11 for programming. I don’t think I’ll lose that ability. It’s a native tongue.

    But when I create tools like this, I cycle through “will this cause my skill to atrophy,” to “do they even matter,” settling on “can you evaluate without creation?”

    When the tool says there’s something wrong? Often correct. When the tool makes an error? It’s usually slips or failures of a global mental model. The internal validation means the mistakes are edge cases not encountered yet. Sometimes it breaks something visible. But like you see above, sometimes it’s just potential.

    I like crafting tight, efficient queries, but I’m also proud of this tool, making something that’s eliminated an entire class of problems. The benefit is not theoretical. I’m not worried about AI taking my job. If all the tasks had been simple enough for an AI to take, then the job wasn’t worth doing in the first place. They aren’t.

    What AI does becomes the floor, and it does it without understanding. Patterns. My job isn’t to craft complex SQL queries. That is an effect, an output, evidence from a mental model.

    I worry people will equate output with judgment.