Latent Self-Expression

The Room and the Record

You can build this today, and plenty of teams already have. Give an agent everything a project has produced: every Slack message, every deck, every status update since kickoff. Put a dashboard and a chatbot on top, so anyone can ask where things actually stand and get the real answer instead of the curated one. Call it a panopticon over the project, because that's what it is, and nobody in the room objects because the pitch sounds like accountability rather than surveillance.

The use case that sells it is greenwashing, in the internal sense: project updates get written for an audience, not for accuracy, and everyone in the building already knows it. So you want a system that reads across the official record and catches the gap between what's being reported and what's real. Something that watches for progress being performed rather than made, the corporate equivalent of a lie detector that only works on memos.

In bounded territory it's genuinely good. It pulls facts across years of history, tracks what people committed to, catches the exact moment a Q1 promise quietly slid into a Q3 maybe. Retrieval was never going to be the hard part.

The conclusions are where it comes apart. They come back consistently, frustratingly off, never wrong in a way you can actually point to and win the argument. The agent reads the content of every message and misses the thing the messages were traveling on. The name that's conspicuously absent from a thread where it should appear. The update that softened every hedge that was live in the draft version an hour earlier. The person who went quiet in the meeting at the exact moment their silence became the loudest thing in the room. None of that survives in a Slack export, because none of it was ever written down in the first place. You can't subpoena a pause.

Two problems inside one quote

The Economist ran a piece in June on AI and tacit knowledge that opened with Polanyi's line: we can know more than we can tell. It used the line to frame a problem about handing expertise to AI agents. The example was a bricklayer who can't explain why he vibrates his hand when he sets a brick into mortar, until someone films him and works out that the motion pushes mortar into the brick's pores and strengthens the bond. He knew more than he could tell. The camera told it for him, eventually, which is a nice enough story if you don't ask it to do more work than it can bear.

I've written before about why Polanyi's argument is harder than that framing makes it sound, about how looking directly at the parts of a skilled process can dissolve the process instead of revealing it. The ASML lithography machines that nobody's managed to replicate off the stolen blueprints are the clearest case I know. The blueprints are complete. They still don't contain the thing, and China has now spent the better part of a decade proving it the expensive way.

But the Economist piece runs two different problems together, and they come apart in a way that actually matters for anyone deciding what to automate next quarter.

The first problem is tacit knowledge: know-how built through iteration, laid down in a body over years of doing. The bricklayer's hand. The forty-year engineer who reads a machine the way you read a face, and can tell you something's off before the dashboard agrees. This kind of knowledge resists language, sure, but it's in there. Given enough observation, in principle, you can get it out. The camera got the bricklayer's.

The second problem is the one that actually breaks the greenwashing system, and it's shaped differently. Some things are communicated only in the act of communicating, and they don't exist anywhere outside that act. Tone. The length of a pause before someone answers. What a person pointedly leaves out of a message where they otherwise say everything else. Whether "I'm fine with that" means agreement or means a quiet decision to bury the thing in three weeks (you know this tone if you've ever managed anyone; it has a very specific flatness to it). These aren't facts that failed to get written down. They have no written form. They happen in the room and they're gone when the room empties, and no amount of retroactive interviewing gets them back, because you can film a bricklayer's hand but you can't film a hesitation that already happened.

That second channel carries real predictive weight, not just color. In 2012 the Journal of Finance published a study by two Duke researchers, William Mayew and Mohan Venkatachalam, who ran vocal emotion software over the audio of 1,647 earnings calls across 691 companies. They weren't reading the transcripts. They were reading the voices, scoring two states in particular: excitement, and the strain of cognitive dissonance, described in the paper as roughly what you hear when an executive puts a positive spin on numbers that don't support it. The vocal signal predicted the next six months of earnings and stock returns, on top of everything the words and the financials already said. The more negative the voice, the worse the future. And the effect was strongest under pointed questioning, exactly when the strain is hardest to fake your way out of.

The human analysts on those same calls mostly missed the negative signal too, which is the detail I keep coming back to. They moved on the excitement and held back on the strain, waiting for numbers to eventually confirm what their own ears were already telling them. The signal was right there in the channel, predictive, and even the professionals whose entire job was to catch it were slow on the uptake. That's the layer organizations actually run on. Who's really behind an initiative versus who put their name on it to avoid a fight later. Which senior person's "let's explore that" means yes and whose means a polite funeral. Where the coalition sits that will quietly kill the thing everyone just nodded along to in the room. It moves through tone and timing and omission, and the written record holds the words while losing exactly the signal that mattered.

Scaffolding versus replacing

There's a distinction here that tends to get collapsed, and the whole question of what these systems are actually good for sits right on top of it.

An agent that hands an expert the right information so the expert can navigate better is one thing. An agent trying to do the navigating itself is a completely different thing wearing the same demo. Scaffold a person with retrieval and synthesis and you get something genuinely powerful: the human reads the channel, the agent holds the context and the memory, and the combination clears a real bar. That version works. It's most of the value actually on the table right now, unglamorous as "the agent fetches things well" sounds next to "the agent manages the org."

The trouble starts once you push past scaffolding into full automation of tasks that are made of the channel, not merely informed by it. Managing people whose interests don't line up. Telling the difference between a project that's healthy and one that's performing health for the steering committee. Working out why something that looks fine on paper is quietly dying underneath it. Channel signal isn't a nice enrichment layer on top of these tasks. It's the substance of them. You cannot do the work without reading it, any more than you can referee a match by reading the box score after it's over.

So the agent does the only thing it can do. It produces output that pattern-matches to what an informed person would say. It writes a credible email. It names plausible risks, in a plausible order, with plausible confidence. Everything it returns is coherent and defensible and reads great in the demo, and it's still missing the part that mattered, because the gap between what it knows and what a person in the room knew was never sitting in the text to begin with. It was in the room. It stayed in the room.

The demo problem

This is why the failure is so hard to catch before you've already committed to it.

When you evaluate an agent for this kind of work, you test the things you can actually see: factual accuracy, coherence, quality of the writing. The agent passes all of them, cleanly. The demo goes well, the decision gets made, and the failure shows up six months later when something political goes sideways in a way the agent's careful, well-footnoted analysis never once gestured at.

I keep finding the same shape in different places. In the gap between an AI passing a benchmark and actually having the capability the benchmark stood for, the benchmark quietly becomes the thing you're optimizing, and the capability it was meant to track stops being specified at all. Same move here. The transcript is a benchmark for the meeting. It's legible, it's complete, it's auditable, and it's missing exactly the part that doesn't reduce to text. When you build a system that reads the record, the record becomes the territory, and the room where the real decision actually happened drops off the map entirely. The proxy is always the part you can measure, which is exactly why the proxy gets automated first, and exactly why what it left out stays invisible right up until it costs you something real.

There's a version of this same gap showing up on the interpretability side of the field too, which suggests it's not a Slack-export-specific quirk but something closer to a structural feature of legibility itself. Natural language autoencoders let you read which features fire inside a model as it processes your input. Genuinely useful, genuinely new. But that only makes the model's own substrate legible. It says nothing about the thing that emerges in the loop between you and the model, the question you'd never have thought to ask without its last reply, the framing its output made available. That part isn't hiding in the weights waiting to be decoded. It doesn't live in a substrate at all. It lives in the interaction, the same way the pause before someone answers lives in the room and nowhere else.

There are models being built to close some of the organizational gap too. Architectures that take in audio directly and read prosody alongside the words. Hume's voice work is aimed squarely at decoding emotional register; OpenAI's audio models hold tone in the representation before transcription flattens it out of existence. Mayew and Venkatachalam needed a piece of Israeli software to score a voice back in 2012, so none of this is alien to machines in principle. Whether it ever reaches organizational subtext, reading a room, sensing a buried coalition, noticing that the project sponsor hasn't said a word in twenty minutes and that this is the actual news of the meeting, is a separate question. Nobody can answer it honestly yet, and I'd be suspicious of anyone who claims they can.

Where this leaves you

If your agent deployments on the human-layer tasks feel slightly off in a way you can't quite locate, this is usually the reason. The system is reading the record faithfully, fully equipped and well prompted, and the meeting still happened in the room.

Call it a placement problem, not a verdict on these systems or a case for human superiority. Point them hard at the work that's made of information: retrieval, synthesis, tracking, analysis, the places where the record actually holds what matters. Move slowly on the work that's made of the channel, because the channel is the one thing the agent structurally can't get to, no matter how much audio you feed it. The people who get the most out of this generation of agents will be the ones who can tell, before they deploy anything, which of those two kinds of work they're actually looking at. Everyone else gets a very confident memo about a meeting that never happened the way the memo says it did.