PP&A Case Study
Designing and Testing a Survey That Benchmarks Generative AI Developer Tooling
How a Fortune 10 technology company readied a developer survey for launch with cognitive testing
Client Situation

A group that builds AI development tooling at a Fortune 10 technology company wanted an internal benchmark of how well it supports developers who take generative AI features from prototype to production. Its product and user experience research leads framed the question around the jobs developers do, from choosing and training a model to getting tooling and process support. They wanted to know how effective developers felt at each step, which steps were red, yellow or green, whether the wider industry had already solved the problems their own teams faced, and how the company stacked up.

The plan was to put the same questions to two groups: the company's own engineers working with its internal tools, and external developers who had launched generative AI features on the major cloud AI platforms. That comparison would hold only if external developers read every question the way the company intended. The draft grew out of internal work, and the team did not know whether its workflow model matched how other organizations build generative AI. It asked PP&A to help shape the survey and to test it with developers before launch.

Our Approach

PP&A scoped the work with the client in December 2024 as two steps: shaping the survey instrument, then cognitive testing with developers in proctored sessions. The instrument follows a developer through five phases of the generative AI lifecycle: exploring viability and planning, experimentation and optimization, evaluation, productionizing, and sustaining and evolving a launched system. Developers rate how well their tooling enables 25 tasks within those phases, from comparing foundation models to obtaining governance approvals, on a five-point scale.

PP&A built two alternative formats of this workflow section for testing. The first asked which phases a developer had worked on and then rated the tasks within each phase. The second asked about involvement task by task, with a short label for every task, across ten screens. At the end of February 2025 PP&A wrote a cognitive testing guide and led sessions with developers from the survey's target audience. Each participant first worked through the survey on screen without help, answering as a real respondent would. PP&A then took them back through each screen to probe vocabulary, concepts and relevance to their role, and asked whether their understanding, or their answer, had changed. The client's own probes set the focus: what "directly involved" should mean, whether respondents would read "effectively" and the question stem as intended, and whether the five phases reflected work outside the company.

Client Results

The client received a launch-ready instrument, finalized in March 2025. It screens for full-time software developers with at least two years in the role who are building generative AI applications and have launched a feature on a major cloud AI platform in the past six months, across four world regions. It also captures each respondent's use case and system inputs and outputs, and takes 15 minutes or less.

Between the tested draft and the final version, the instrument changed on each point the testing probed. The phases gained plain-language definitions tied to concrete activities, such as building a proof of concept. The broad stem asking how well "your organization and its tooling" equipped a developer became a question about how well the tooling and infrastructure of the respondent's primary cloud platform enabled each task. A new answer option separates tasks a developer does not do from tasks done outside the platform, and every phase now ends with an optional open question on where the tooling fell short or where developers prefer other solutions. An internal term for compute budgets gave way to a neutral one, and two optimization tasks now name latency and cost rather than abstract system constraints. The final version kept the phase-first format.

The instrument went on to run as a tracking study, with a first wave in March 2025 and a second in the third quarter of 2025 on the same five-phase structure. Fielding and analysis sat outside PP&A's scope. The company gained a repeatable measure of how well cloud AI tooling supports each stage of generative AI development.

  • --
Share this case study: Share on LinkedIn
Ready to Elevate Your Business?

Our expert teams are committed to helping you navigate market complexities and transform challenges into substantial opportunities. Explore how our specialized services can drive your professional success by connecting with us today.

Contact Us
About PP&A

At PP&A, we are dedicated to empowering businesses to excel through strategic innovation. We integrate deep industry insights with comprehensive analytics to develop bespoke solutions that propel growth and enhance performance. Our services encompass a broad spectrum, including:

Executive Training
Ready to Implement AI in Your Organization?

Join our Generative AI Executive Workshop and gain the strategic insights and hands-on skills to lead AI-driven transformation. Starting at $2,750 per participant.

Related Case Studies