ashaykubal.com
August 3, 2026

Welcome to Enterprise AI Theater

Welcome to Enterprise AI Theater: the performance of transformation, staged for a budget committee

If your head of product and your head of AI have decided that the way to accelerate your business workflows is a chatbot, you need a new head of product and a new head of AI.

Bold statement to make, but in this post I will defend that. It is not because chatbots are inherently useless. It is because handing a pretty chatbot that answers some questions to the business sponsors or users is what a demo looks like. Your product lead watched a model summarise a document, or pull some information and present it, then they demoed it to you, felt the room go quiet and read that silence as a roadmap. If you are the business sponsor or the user, you are going to pay for the difference much later, in a line item nobody wants to own. Figuring out the distance between a demo and what it takes to qualify a release candidate, then outlining that journey to you so that you know what you are getting into, is the entire job: a job your product and AI leads did not do.

This is not an argument against AI. I build with these tools every day and I would not go back. It is an argument against theater: the performance of transformation, staged for a budget committee, by people who have confused the part that photographs well with the part that does the work.

The house lights are up

Let us start with the number everyone quotes, something like "88 percent of organisations are using AI in at least one function." You have seen it, or one of its cousins, in every deck this year.

The United States Census Bureau measures the same thing on a probability sample of actual firms and gets 17 to 20 percent.1 Eurostat, measuring European enterprises, gets 20.0 percent.2 The gap is partly definitional, and I am not accusing anyone of lying. But the most repeated statistic in enterprise AI is four to five times what the government's statisticians find, and the difference is who agreed to answer the survey. That is the whole subject of this article, demonstrated on its own most-cited number, before we have started.

Now, measurement. The Federal Reserve Bank of Atlanta surveyed 748 chief financial officers. They reported AI-attributable productivity growth of 1.8 percent for 2025. Their own reported revenue and employment figures imply 0.6 percent.3 Same executives, same year, same firms, a gap of three times between what they believe and what their numbers show.

It repeats at the individual level. METR ran a randomised trial on experienced open-source developers working in their own repositories. They were 19 percent slower with AI assistance. They believed they had been 20 percent faster.4

Three populations, three methods, one pattern.

The perception of what AI can achieve is bright, but the measurement of what it achieves is thin.

A word about who is telling you this

You will notice I have not cited a management consultancy, and I want to be direct about why.

The most repeated framework in enterprise AI says the work is 10 percent algorithms, 20 percent technology and data, and 70 percent people, process and cultural transformation. It is a good insight. It is also, word for word, the scope of a transformation engagement, published by the firms that sell transformation engagements and used by the executives who hire them to secure budgets for said transformation theater. Everyone in this industry knows what it means when the consultants arrive to tell you your AI problem is a culture problem. We say it to each other in the lift. We do not say it in the steering committee.

The underlying finding is real, and it is older and better than the framework. Back in 2002, economists studying the last big wave of enterprise technology found something strange: a dollar a company spent on computers showed up as roughly twelve dollars of company value, while a dollar spent on ordinary equipment was worth about a dollar and a half.5 The extra ten dollars was never the computers. It was everything the computers forced the company to change: the processes redesigned, the roles rebuilt, the new ways of working that never appear on a balance sheet. Peer-reviewed, replicated, and about enterprise software rather than AI. So the claim that the organisation matters more than the technology is a twenty-five-year-old economics result with a ratio near eight to one, and the version you were sold is a rebrand with invented weights.

The symptoms of the theater have names. "Workslop," coined by researchers at Harvard Business Review with Stanford's Social Media Lab, is AI output "that masquerades as good work but lacks the substance."6 Gartner calls the vendor version "agent washing."7 And the executives presenting the AI strategy are the group most likely to be using unapproved tools regularly.8 The receipts are already public, if you want them: a Canadian tribunal held Air Canada to the refund policy its chatbot invented, and the SEC's own chair coined "AI washing" while charging advisers for selling AI they were not using.9, 10

Your chatbot is answering a 2024 question

For thirty years the deal was that work lives in applications and you go to it. Open the customer relationship management system, then the reporting tool, then the spreadsheet, then the inbox, and carry the context between them in your head.

That deal has ended. Not everyone has noticed, and the first section of this article already showed you how wide the gap between what happened and what people believe can run.

Look at what shipped inside twelve months. Anthropic's Claude Cowork, Microsoft's Copilot Cowork, OpenAI's ChatGPT Work, Google's Gemini Enterprise, Amazon's Quick Suite: every major vendor now sells the same shape, one surface where you actually work, into which applications, dashboards and agents arrive.11 Open source kept pace: OpenWork is an open rebuild of the same idea, and Andrew Ng released OpenWorker in July, a local agent that hands you finished documents, sent messages and updated calendar entries rather than chat.12 Nous Research's Hermes Agent runs the pattern on the desktop with Nvidia's enterprise stack integrated behind it.13 Even the picks-and-shovels layer has reorganised around the shape: Nvidia launched an agent development platform with seventeen enterprise software firms building on it, and LangChain's tooling has grown from observability into a full operations stack for agent fleets.14

The chatbot was one surface bolted beside many applications. Fine for its time. The funnel has now inverted: the surface is the application, with the systems you depend on arriving into it as tools, rendered forms and finished work.15

I can hear the objection, because I opened this article by firing people for proposing exactly this. Look again at what got fired. That chatbot was a widget at the side of each application, one more surface on the pile, answering questions about a world it could not act in. Fifteen applications became fifteen applications plus fifteen chat windows. What I am describing runs the other way. The applications move inside the surface, records show up already updated, and the pile of things you log into starts shrinking instead of growing. There is a catch, and it is the one that matters: you cannot buy your way across that distinction. Take any of the surfaces I listed above, connect nothing to it, define nothing underneath it, and congratulations, you have rebuilt the old chatbot on a bigger stage. The surface only gives you the venue. What has to be true before real work happens in it is the rest of this article.

The reason this shift will stick is the plumbing underneath it. The Model Context Protocol, the connector standard the industry converged on, went stateless in July: connectors now behave the way websites and APIs have behaved for decades, plain request and response, which is what lets any surface talk to any system without a standing relationship between them.16 And the protocol is nobody's moat: it now lives at the Agentic AI Foundation under the Linux Foundation, with more than 5,800 servers shipped against it.17 The obvious objection, that this trades application lock-in for surface lock-in, dies on that fact. The vendors will try to charge rent anyway. Switching surfaces is now cheaper than switching applications ever was.

The incumbents can read the funnel as well as anyone. Salesforce announced Headless 360 on the premise that the platform is reachable through interfaces and protocol tools rather than a browser, and its co-founder asked the question out loud: "Why should you ever log into Salesforce again?"18 In financial services the movement is further along than most. S&P Global built an official MCP server with Anthropic, putting Capital IQ financials and earnings transcripts directly into the agent surface.19 LSEG shipped MCP connectivity across seven platforms in twelve months.20 And Anthropic now ships plugins scoped to business groups: wealth management, equity research, investment banking, private equity.21 A model vendor does not build a wealth management plugin unless it expects wealth managers to be doing their day's work inside its surface.

I have also run the experiment at personal scale. I built a working connector for Redtail, a customer relationship management product used by financial advisers, by reverse engineering its traffic in browser developer tools before I had any access to its official documentation; Google ships a developer-tools protocol server that exposes exactly that inspection to agents, so the route is tooled rather than heroic.22 Then I built the other half: create-and-edit forms that render inside the agent, styled like the product they came from, so the record arrives where I already am. No swivel-chair. If one person can do that with tooling given away free, the hard part of this transition is somewhere else.

For the firms that will build rather than buy, the serving layer is also free. Amazon, Google and Microsoft have open-sourced their agent frameworks, and Amazon's Bedrock AgentCore Gateway, generally available since October 2025, composes many backing connectors behind a single front door with one consolidated tool list and a semantic tool search across everything registered.23, 24 That is the shape of an enterprise workspace: one entrance, many systems behind it, discovery that works.

I should be honest about where this still falls short. The in-surface rendering that makes it all seamless is young, experimental by its own documentation, and you will not find a catalogue of ready-made agent apps to install.25 I would not wait for one. The standard fixes how apps render inside the surface and leaves the apps themselves to you, which is how it should be, because your forms need to carry your design system, your terminology and your entitlements. The other caveat is adoption. Everything I have described has shipped, but shipping is the vendors' half of the story, and the first section of this article should tell you how much daylight to expect between what ships and what gets used.

Back to the point. This is what the modern work operating system looks like: One front door. The systems you depend on reachable behind it as tools. Forms rendering in front of you, in your design language, when the work calls for them. Records updating where you already are, deliverables arriving finished, agents carrying the context between systems so you no longer have to. Every piece of that has shipped.

I will do a deep dive into the modern enterprise AI architecture over two posts after this one. Which leaves this article to tackle planning and scoping.

The unified work operating system is an outcome, not the strategy.

What stands between you and that outcome has nothing to do with interfaces or tools.

The only question that matters

This is the part that is hard.

Enterprise AI should be outcome-tagged, and the sequence does not bend. Sit with the business and establish the outcome. Define what done looks like and how you will know. Then, last, specify the minimum AI implementation that gets there.

That last step is where the discipline lives, and the evidence for it is unglamorous. Flyvbjerg's database of 1,355 programmes finds each additional year of duration adds 4.2 percentage points of cost risk.26 Longer is worse, reliably, across every project class anyone has measured. Fewer things, defined properly, shipped sooner.

Which means you have to be willing to reject one. Outcome-tagging becomes a discipline only when some outcomes fail the test, and the test is rarely whether the model could do it. The real price of an outcome is what it would cost you in meaning: how much of your data would have to be made comprehensible before the thing could work, who would have to sit down and agree what the fields mean, and who owns that agreement afterwards. An outcome that requires three desks to settle the definition of revenue is an expensive outcome, whatever the demo looked like. That price is knowable before you start. Working it out, and saying no to the ones that fail, is what separates a portfolio from a wish list.

Write the evaluation before you build. It tells you whether you got there, and it turns out to do a second job.

And your data is a piping hot mess

Snark back on, because this one earns it.

Chatbots are magnificent on clean data. Yours is not clean. Your contacts live in five systems that disagree about the client's legal name, or worse, some copies may not be tagged to one at all. "Revenue" means three different things depending on whose compensation depends on the answer. Point a language model at that and you have automated the confusion and given it a confident voice, which is worse than the confusion, because the confusion at least sounded uncertain.

Here is what the numbers actually say. Anthropic published how it runs its own internal analytics, and the figures are worth your attention because they argue against the interest of a company selling models. Without a curated context layer, its agents answered at 21 percent accuracy. With one, 95 percent and above, regularly 99 in some domains.27 Independently, the Spider 2.0 academic benchmark measured frontier models at 21.3 percent on realistic enterprise query workloads.28 An internal practitioner measurement and a peer-reviewed benchmark, different methods and opposite motives, landing a third of a point apart. Point a good model at an enterprise data estate with no curated context and roughly one answer in five is right.

Three findings underneath that headline matter more than the headline. Giving the agent direct access to thousands of prior SQL queries improved accuracy by under one point, so more data changed nothing. Auto-generating the definitions failed, and their recommendation is exact: generate the documentation with the model, but have "a human own the definition." And without maintenance it rots, drifting from 95 percent to 65 percent within a month.27

The reason is documented. Ninety-six percent of enterprise questions require knowledge present in neither the question nor the schema.29 The missing thing is meaning, and meaning does not live in your warehouse. Those findings are facts about how these systems work rather than facts about how tidy your tables are.

More context beats more data. Curation beats generation. Everything you build decays unless somebody owns it.

You cannot arrive at any of those by procurement, and a leader who does not hold them will keep funding the wrong half of the problem.

The instinct is to fix everything first. A two-year enterprise data programme, after which the AI will work. That programme is theater too, and Flyvbjerg already told you how it ends.

So build the meaning, and only as much as you need. Not an enterprise ontology. Just enough for the outcome in front of you: the entities and relationships that outcome touches, the taxonomy, the handful of business questions it must answer, the catalog entries and the annotations that say what a field actually means. Then a guardrail, in the same breath, because "just enough" without a fallback is an excuse: when a question falls outside the scoped meaning, the system says so instead of guessing. A plausible but wrong revenue number in front of a salesperson is the worst failure mode available.

How much is enough? The evaluation you already wrote. The scope is sufficient when the system answers the outcome's question set correctly and detects when a question sits outside it. Both halves are measurable, so this is an eval rather than an opinion. Any product manager will recognise the shape of that decision, because shipping was never about a bug-free product. It was about whether the thing does its job, and whether the known defects are ones you can live with, fix later, or ignore. Same test, applied to meaning.

The firms whose entire business is data have known this for years. Bloomberg published its practice for managing annotation in 2020, and the sentence to sit with is this one: "All annotation projects, regardless of their downstream use, will require some budget and resource allocation in perpetuity for monitoring and quality assurance."30 Elsewhere in the same paper: "There is no static data: ongoing annotation is necessary to adapt downstream models/solutions as data drift occurs."30 Bloomberg wrote that six years before anyone measured a context layer decaying from 95 to 65 percent in a month, and its answer was permanent budget. Its annotators are described as a dedicated team of financial experts, curating across forty years.31 LSEG attacks the same problem from the other end, productising meaning rather than staffing it: an ontology-backed semantic layer with persistent identifiers.32 The seven MCP connectors it shipped earlier in this article are the visible end of that work; a connector to your data is only as useful as the metadata enrichment already sitting behind it, and the enrichment is the part that took years. These are the firms every buy-side and sell-side desk depends on, and none of them treats meaning as a project that finishes.

There is a reason those programmes are not the theater I called out a few paragraphs ago. Bloomberg and LSEG sell data. Meaning is the product, so a multi-decade programme across the whole estate is scoped to the outcome, because the outcome is the catalogue itself. Your firm sells something else. Your outcomes are narrower, they are specific to you, and they will not look much like your peers' either. So take the lesson from the data firms rather than their blueprint: own meaning permanently, the way they do, but build it only for the outcomes you actually have. The full-estate version makes sense when the estate is what you sell.

The counterargument is real and I am not going to duck it. A semantic layer scoped to one outcome does not reconcile your three definitions of revenue. It ratifies one and leaves two. Repeat that per outcome and you rebuild the fragmentation you were fleeing, one narrow layer at a time. The industry evidently agrees this is the risk, since Snowflake, Salesforce, dbt and Databricks have donated an interoperable semantic model standard to the Apache Software Foundation.33 Which is why governance starts at the first outcome rather than the fifth: a named owner per definition, a reviewable change history, versioned crosswalks where two desks disagree, and an adjudication path when they cannot.

Even the chief financial officer of the largest US bank says the unglamorous half is unfinished. At JPMorgan Chase's 2024 Investor Day: "70% of our data has been landed in the cloud, but we should caveat that landing is not the same as having it be modernized and usable for AI and ML, so there's work to do there still."34

The seventy percent nobody wants to fund

Now the part where the argument closes on itself.

Who writes that meaning down? Not a vendor, and not a model. Anthropic already tried letting the model do it and got something worse than a smaller human-curated set.27 The business sponsor knows what the measure is supposed to mean. The users know which questions actually get asked. Engineering knows what the system can answer and where the joins break. Compliance knows which answers carry obligations.

The context layer is an artefact that business, users, engineering and compliance produce together, and there is no other way to produce it. The cross-functional argument stops being a virtue and becomes a dependency.

You cannot buy this layer, generate it, or scrape it from your query logs. It exists only if those people build it, and it stays true only if someone owns it after launch. That is the seventy percent, and Bloomberg's word for the budget line is "in perpetuity."

Deciding that room exists, funding it, and holding it together past launch is a leadership decision, and no amount of product management substitutes for it. A product manager can name who belongs there and why, make the vision legible enough that people want to come, and stay accountable for what the group produces. Driving the change sits upstairs. The most reliable way this fails is a product team trying to carry an organisational change on borrowed authority.

Which leaves the question in the heading. Who pays for it?

The answer has been sitting in the first half of this article. Look at what the present arrangement already costs you. Data leaves the customer relationship management system and the trading systems, moves through a pipeline, gets layered and refined, lands somewhere, and then gets piped out again into another application with another backend and another interface to keep alive. Each of those applications carries its own integration surface, its own serving layer, its own team. Most of that machinery exists so that data can be assembled in a place a person will then travel to.

Put the meaning in one place and let the work come to people instead, and a good part of that estate becomes a decommissioning candidate: the duplicated pipelines, the bespoke backends, the surfaces nobody opens once the record arrives where they already are. The people, the infrastructure and the time released by retiring it is what funds the seventy percent. The context layer stops being a cost centre and becomes the thing that pays for the change.

One condition decides whether that is a business case or a fantasy. The decommissioning has to be planned, funded and enforced from the start, with dates. Our industry's record on switching things off is poor. Build the new thing, keep everything else running, and you have added cost rather than moved it.

That business case is also the most useful thing a product manager can hand a leadership team, which is the honest answer to the sphere-of-influence problem above. You will not win the culture argument by making it. You might win it by showing what the current arrangement costs.

Ethan Mollick, professor at the Wharton School and author of Co-Intelligence, put the mechanism plainly: "the moderating factor is organizational structure. It's not individual ability or even AI ability."35

And it works when it is done. Morgan Stanley reached 98 percent adoption across its advisor teams, and got there by defining a small number of outcomes, building evaluations that advisors themselves graded before rollout, grounding the system in proprietary content, and running real change management.36 Worth saying honestly: that figure measures adoption, not return. It is the best kind of adoption number, earned rather than provisioned, and it is still an adoption number.

The chatbot was never the strategy

Which brings me to the thing I found most clarifying, and it is an absence.

I checked the four institutions most capable of measuring an AI return, most incentivised to disclose one, and most scrutinised by investors: Bloomberg, JPMorgan Chase, Morgan Stanley and BlackRock. Between them they have disclosed one target and one refusal to attribute. JPMorgan named $1.5 billion at its 2023 Investor Day, as a target.37 By February 2026 its chief financial officer would say only "$600 million in efficiencies, some of which are AI-related."38 Morgan Stanley's last two shareholder letters contain no AI figure at all.

Firms have legitimate reasons to avoid breaking out attribution, and absence of disclosure is not absence of return. But the pattern holds from the survey data down to the individual developer: what people believe about this technology and what anyone can measure have not yet met.

So do the work in the right order. Define the outcome with the people who own it. Write the evaluation that proves it. Build just enough meaning over your data for that outcome, with a guardrail for its edges and an owner for its upkeep. Redesign the workflow around the result. Then, at the end, choose the interface. If you have done all of that and the right answer turns out to be a chatbot, build the chatbot, and it will be a good one, because you will have arrived at it instead of starting there.

Successful enterprise AI takes three things. Enough conceptual understanding of how these systems actually work to make sound decisions about them. System design that puts them somewhere useful. And the product judgement to pick an outcome worth the trouble, and to know who has to be in the room when you do. Only one of the three is about AI.

Sources

  1. [Tier 1] U.S. Census Bureau, Business Trends and Outlook Survey: "Large Firms With at Least 20 Employees Biggest AI Users," America Counts, May 2026. Biweekly probability sample, nationally representative of nonfarm businesses; 17-20 percent AI use December 2025 through May 2026; 37 percent among firms with 250+ employees. census.gov
  2. [Tier 1] Eurostat, "20% of EU enterprises use AI technologies," 2025-12-11. 2025 EU survey on ICT usage and e-commerce; enterprises with 10+ employees. ec.europa.eu
  3. [Tier 1] S. Baslandze et al., "Artificial Intelligence, Productivity, and the Workforce: Evidence from Corporate Executives," Federal Reserve Bank of Atlanta Working Paper 2026-4, March 2026. The CFO Survey (joint with Duke University and the Richmond Fed); 748 senior financial executives, two waves, November 2025 to January 2026. atlantafed.org
  4. [Tier 2] J. Becker, N. Rush, E. Barnes and D. Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," METR, arXiv:2507.09089, July 2025. Randomised: 246 real tasks, 16 experienced open-source developers in their own mature repositories. METR is an AI-safety research nonprofit; sells no product or advisory service. Preprint, not peer-reviewed, hence Tier 2. arxiv.org/abs/2507.09089
  5. [Tier 1] E. Brynjolfsson, L. Hitt and S. Yang, "Intangible Assets: Computers and Organizational Capital," Brookings Papers on Economic Activity, 2002. See also Bresnahan, Brynjolfsson and Hitt, Quarterly Journal of Economics 117(1), 2002. Studies enterprise systems, not AI. Reported association: roughly $12 of market value per dollar of computer capital, against about $1.47 per dollar of plant and equipment; the excess is attributed to organisational capital. brookings.edu
  6. [Tier 2] Harvard Business Review with BetterUp Labs and Stanford Social Media Lab, "Workslop," 2025-09-22. Term only; the cost estimates are excluded here as the research is commercially interested.
  7. [Tier 1] Gartner, 2025-06-25.
  8. [Tier 2] UpGuard shadow AI research, 2025-11. Executives lead on regular and frequent use; overall rates are highest among mid-level staff.
  9. [Tier 1] Moffatt v. Air Canada, 2024 BCCRT 149, British Columbia Civil Resolution Tribunal, 2024-02-14.
  10. [Tier 1] U.S. Securities and Exchange Commission, press release 2024-36, 2024-03-18.
  11. [Tier 1] OpenAI, "Introducing workspace agents in ChatGPT" (research preview 2026-04-22) and ChatGPT Work launch, 2026-07-09. Microsoft 2026 Work Trend Index and Microsoft 365 Copilot update (Copilot Cowork, Agent 365, Copilot Chat), 2026-05-05, microsoft.com. Google Gemini Enterprise and Amazon Quick Suite product documentation. Anthropic Claude Cowork: first-hand use; this article was researched and drafted in it. openai.com
  12. [Tier 1] A. Ng, OpenWorker announcement, 2026-07-23, MIT licence, local-first, typed approval classes; repository on Ng's personal GitHub. Corroboration: MarkTechPost, 2026-07-23. [Tier 3] github.com/different-ai/openwork, open-source agent workbench built on OpenCode.
  13. [Tier 1] NVIDIA blog, "Hermes Unlocks Self-Improving AI Agents, Powered by NVIDIA RTX PCs and DGX Spark," blogs.nvidia.com; NemoClaw runtime and Nemotron 3 Ultra integration. nvidia.com
  14. [Tier 1] NVIDIA Newsroom, "Enterprise Software Leaders Build AI Agents With NVIDIA," 2026-06-01; seventeen named adopters including Adobe, Salesforce, SAP, ServiceNow. [Tier 2] LangChain Interrupt 2026 releases (LangSmith Engine, Managed Deep Agents, Fleet). nvidianews.nvidia.com
  15. [Tier 1] MCP SEP-1865, MCP Apps, status Final, Extensions Track, specification 2026-01-26.
  16. [Tier 1] modelcontextprotocol.io/specification/versioning, current revision 2026-07-28; PyPI mcp 2.0.0, published 2026-07-28. Anthropic's description of the change: "from a bidirectional stateful protocol to a request/response model," claude.com, 2026-07-28. modelcontextprotocol.io
  17. [Tier 1] Anthropic, "Donating the Model Context Protocol and establishing the Agentic AI Foundation." Foundation under the Linux Foundation, co-founded with Block and OpenAI; Bloomberg among supporters. Server and SDK ecosystem figures as of mid-2026. anthropic.com
  18. [Tier 1] Salesforce Newsroom, Headless 360, 2026-04-15. [Tier 2] diginomica, Parker Harris, 2026-04-01. The named MCP Server component was in beta as of 2026-07-14.
  19. [Tier 1] S&P Global press release, "S&P Global and Anthropic Announce Integration of S&P Global's Trusted Financial Data into Claude," 2025-07-15. Kensho LLM-ready API MCP server; S&P Capital IQ Financials and earnings call transcripts. press.spglobal.com
  20. [Tier 1] lseg.com AI finance solutions pages, MCP integrations across seven platforms (Databricks, Microsoft, Anthropic, OpenAI, Google Cloud, Amazon, Model ML), 2025-08 through 2026-07, retrieved 2026-07-28. The shipping record is cited as fact; LSEG's claims about the value of AI-ready content are commercially interested and are not relied on.
  21. [Tier 1] Anthropic, "Claude for Financial Services"; plugin pages at claude.com/plugins (wealth management, equity research, investment banking, private equity), Claude Enterprise. [Tier 2] Finextra, "Anthropic launches financial services plugins for Claude Cowork," 2026. anthropic.com
  22. [Tier 1] github.com/GoogleChrome/chrome-devtools-mcp. github.com
  23. [Tier 1] AWS Strands Agents (Apache 2.0, 2025-05-16); Google Agent Development Kit (Apache 2.0, 2025-04-09); Microsoft Agent Framework (MIT).
  24. [Tier 1] AWS documentation, Bedrock AgentCore Gateway, generally available 2025-10-13.
  25. [Tier 1] Goose documentation describes its MCP Apps support as experimental and in active development.
  26. [Tier 1] A. Budzier and B. Flyvbjerg, "Overspend? Late? Failure? What the Data Say About IT Project Risk in the Public Sector," BT Centre for Major Programme Management, Said Business School, Oxford, December 2012. arXiv:1304.4525. n=1,355 public-sector IT projects; each additional year of duration adds 4.2 percentage points of average cost risk; projects beyond 24 months disproportionately represented among fat-tail overruns. arxiv.org/abs/1304.4525
  27. [Tier 1, commercially interested, admitted on adverse interest] "How Anthropic enables self-service data analytics with Claude," claude.com, 2026-06-03. Cited because its findings argue against the publisher's own product.
  28. [Tier 1] Spider 2.0, ICLR 2025 (oral).
  29. [Tier 1] EntSQL, HKUST with Alibaba, 2026-07.
  30. [Tier 1] T. Tseng, A. Stent and D. Maida, "Best Practices for Managing Data Annotation Projects," Bloomberg, arXiv:2009.11654, September 2020. Internal practice documentation, not marketing. arxiv.org/abs/2009.11654
  31. [Tier 1] BloombergGPT, arXiv:2303.17564, 2023. arxiv.org/abs/2303.17564
  32. [Tier 1] LSEG Intelligent Tagging and PermID product documentation.
  33. [Tier 1] Apache Software Foundation, donated interoperable semantic model standard (Snowflake, Salesforce, dbt, Databricks).
  34. [Tier 1] JPMorgan Chase Investor Day, 2024-05-20, Jeremy Barnum.
  35. [Tier 2] Ethan Mollick, interview, 2026-03.
  36. [Tier 1] Morgan Stanley press release, 2024-06-26. 98 percent of Financial Advisor teams.
  37. [Tier 1] JPMorgan Chase Investor Day, 2023, Lori Beer. Stated as a target.
  38. [Tier 1] JPMorgan Chase, chief financial officer remarks, February 2026.
← Back to writing
© 2026 Ashay Kubal New York · UTC−5 personal lab