Every software services firm now describes itself as AI-native, agentic, or AI-first. The label has become table stakes in a sales deck, which means it has stopped carrying information. For a buyer trying to choose a partner, the claim is worse than useless, because it makes genuinely different companies sound identical.
Gartner has quantified the gap between claim and reality. It estimates that of the thousands of vendors marketing agentic AI, only about 130 are real, with the rest engaged in what it calls agent washing: rebranding existing assistants, robotic process automation, and chatbots as agentic without the underlying capability. Gartner also predicts that more than 40 percent of agentic AI projects will be canceled by the end of 2027, largely on escalating costs, unclear value, and inadequate controls. A meaningful share of those failures will trace back to a procurement decision made on marketing language rather than evidence.
This is a framework for making that decision on evidence. It covers what AI-native actually means, the criteria to score a partner against, the questions that expose agent washing, and the delivery models that determine who really owns the result.
AI-Native vs. AI-Washed: The Distinction That Matters
The difference is architectural, not cosmetic. An AI-native software engineering partner is built around agents and a delivery platform from the ground up, with engineers, processes, and commercial models designed for a world where agents do much of the execution. An AI-washed one is a traditional services firm that added an AI practice, resells third-party copilots, and runs the same labor-based delivery underneath a new label.
The distinction matters because it determines where the value of AI actually lands. In an AI-native model, the productivity agents create is engineered into the delivery and can be passed to the client as faster, cheaper, outcome-priced work. In an AI-washed model, the same tools are bolted onto a headcount-and-hours business that has no structural reason to pass the savings on, and often a reason not to. The buyer pays for AI and receives staff augmentation with better tooling.
The tell is often in the demo itself. An AI-washed pitch shows an impressive isolated capability, a copilot completing code, a chatbot answering a question, and leaves the integration, governance, and production hardening as an exercise for later. An AI-native pitch shows the unglamorous parts: how agents are orchestrated across the lifecycle, how their actions are logged and reviewed, how the whole thing behaves inside a regulated client’s controls. The first is selling a feature. The second is describing a delivery system.
Telling the two apart is the entire job of the evaluation.
The Evaluation Scorecard
Score a prospective partner against six criteria. Genuine AI-native partners clear most or all of them with verifiable evidence. Agent washing tends to fail the same ones, in the same ways.
| Evaluation criterion | What a genuine AI-native partner shows | Agent-washing red flag |
| Proprietary platform | Owns and runs its own agent and orchestration platform across the lifecycle | Resells third-party copilots with a services wrapper and no platform of its own |
| Production proof at scale | Agents running in production, at volume, in regulated environments, with named deployments and metrics | Demos, proofs of concept, and pilots; no verifiable production evidence |
| Platform and agent maturity | Orchestration across the full SDLC, with governance, audit trails, and human-in-the-loop controls built in | Single-point tools with no orchestration, no audit trail, no coordination story |
| Engineering depth | A deep engineering bench that directs agents and owns the architectural judgment | Tool licenses plus generalist staff, with no real engineering ownership |
| Governance and compliance | A track record delivering in regulated industries, with controls engineered in | Governance described as “on the roadmap” or applied after the fact |
| Commercial model | Willing to price to outcomes and put fee at risk against the result | Only time-and-materials or per-seat licensing, with all risk on the buyer |
The scorecard is deliberately weighted toward evidence over description. Any vendor can assert governance or claim production scale. The evaluation is whether they can show it: a named platform, a named deployment, a metric a reference will confirm.
The Questions That Expose Agent Washing
A short set of direct questions separates the real from the rebranded faster than any capabilities deck. Ask them, and listen for whether the answer is specific or evasive.
Ask for a named production deployment, with the agent count and uptime. A genuine partner can point to agents running in a real client environment and describe what they do. Agent washing retreats to pilots and hypotheticals.
Ask what is proprietary versus resold. If the “platform” turns out to be a thin layer over someone else’s model or copilot, the partner is a reseller, and you are paying a margin for integration you could buy directly.
Ask how agent actions are governed and audited. A real answer describes human-in-the-loop checkpoints on consequential actions, audit trails, and controls that run inside the client’s compliance rules. A weak answer describes intentions.
Ask whether they will commit to an outcome. Willingness to price to a measurable result, and to carry some of the risk of reaching it, is the strongest single signal that a partner believes its own delivery claims. A firm confident in its platform will stand behind an outcome. A firm selling hours will insist on selling hours.
Ask how much of their own delivery runs on the platform. A partner that has re-engineered its own engineering around agents can say so concretely. One that sells agentic transformation while delivering the old way cannot.
Ask about certified process maturity. An independent appraisal such as CMMI, for both development and services, signals that the delivery discipline behind the agents is real and audited, not asserted.
Delivery Models: What You Are Actually Buying
Two partners can quote the same work under very different commercial structures, and the structure determines who owns the result. Understanding which model you are being sold is as important as evaluating the technology.
| Delivery model | What you are buying | Who owns the outcome | Fit with AI-native delivery |
| Staff augmentation | Bodies and hours to extend your team | You do | Low. AI productivity accrues to the vendor’s billable hours, not your result |
| Managed capacity | A managed team against a defined scope | Shared, often ambiguous | Medium. Better aligned, but frequently still priced on effort |
| Outcome pods | A committed outcome delivered by a platform and team | The partner | High. Matches AI-native, outcome-based delivery and puts the result on the provider |
The pattern is consistent. The further a model moves from paying for effort toward paying for a result, the more the partner’s incentives align with using AI to deliver faster and better rather than to bill more. A partner that is genuinely AI-native will be comfortable moving toward outcome-based structures, because its economics depend on them. A partner that is not will resist, because outcome pricing exposes the gap between its claims and its delivery.
Reading the Whole Picture
No single criterion is decisive on its own. A firm can own a real platform and still lack the engineering depth to deploy it well, or have strong governance and no willingness to commit to outcomes. The framework works as a composite: score the six criteria, ask the questions, identify the delivery model, and look at whether the answers cohere into a partner that is genuinely built for agentic delivery or one assembling the vocabulary of it.
The 40 percent of agentic projects Gartner expects to fail will not fail because agentic AI does not work. Many will fail because the buyer could not tell a genuine platform from a rebranded tool at selection time, and discovered the difference only in production. The framework exists to move that discovery to the evaluation, where it is cheap, rather than the deployment, where it is not.
The Trap the Framework Is Built to Avoid
The most expensive mistake in this market is evaluating a partner on a proof of concept and discovering the truth in production. A curated pilot, run by the vendor’s best people on a friendly use case, is designed to succeed. It reveals almost nothing about whether the partner can deploy at scale, inside your compliance regime, against your messy data and legacy systems.
This is the capability-deployment gap that sits underneath Gartner’s cancellation forecast: the distance between a demo that impresses and a system that holds up when it is integrated, governed, and accountable. The defense is to weight the evaluation toward evidence of production delivery in conditions like yours, and to treat a polished pilot as necessary but not remotely sufficient. Ask what has run in production, for how long, at what scale, and under whose governance. A partner that can only show you the pilot is telling you where its capability ends.
Applying the Framework: What Genuine AI-Native Looks Like
Held against these criteria, Ascendion is a useful reference point for what a genuine AI-native software engineering partner looks like, because the evidence is verifiable rather than asserted.
On platform, it owns and runs AAVA™, its proprietary agentic AI platform, across the full software development lifecycle rather than reselling someone else’s tools. On production proof, it operates more than 10,000 production agents inside Fortune 500 environments, with named outcomes in regulated banking and healthcare deployments. On maturity and governance, AAVA provides orchestration, human-in-the-loop controls, and audit trails, backed by an independent CMMI Level 5 appraisal for both Development and Services. On engineering depth, more than 11,000 engineers direct the agents and own the judgment. On commercial model, its AI Commercials option prices to outcomes rather than hours, which is the firm putting its fee behind the result. And it was built AI-native from the start rather than adding an AI practice to a legacy services business, a distinction reflected in its recognition as a Market Leader in HFS Horizons: Agentic Services, 2026.
The point is not that a buyer should choose Ascendion. It is that these are the kinds of specific, checkable proof points a genuine AI-native partner can produce, and that any vendor unable to produce their equivalent should be evaluated accordingly. For related decisions, see our work on outcomes-based commercial models, which examines the pricing structures referenced here, and on what AI agent governance should look like.
Evaluating AI-native partners for an enterprise program? See how Ascendion delivers AI-native software engineering in production →
Ascendion is the AI-native disruptor reinventing how global enterprises build software for impact. Its engineering teams, powered by AAVA, the company’s proprietary agentic AI platform, deliver measurable business outcomes: accelerating growth, unlocking capital, and de-risking transformation. With 11,000+ engineering professionals and 10,000+ AI agents working across 12 countries, Ascendion delivers the promise of AI to more than a third of the Fortune 500. Learn more at https://www.ascendion.com.
Engineering to the Power of AI™, AAVA™, and Engineering to Elevate Life™ are trademarks or service marks of Ascendion®. AAVA™ is pending registration. Unauthorized use is strictly prohibited.