Answer capsule
Valence reports that AI-coaching power users in its study had 31% higher odds of moving up a performance band across a full review cycle. That is a provider-reported association, not a 31-percentage-point gain or proof that coaching caused the movement. A buyer should examine who became a power user, how engagement and performance were measured, which alternative explanations were addressed, and whether a prospective evaluation can separate selection from effect.
What the source establishes
- Valence's August 18, 2026 provider article says its first Valence Research Initiative study followed 11,000 employees at one large technology company through a full performance review cycle.
- The provider reports that Nadia power users had 31% higher odds of moving up a performance band; this is an odds statement, not a 31-percentage-point increase in probability.
- Valence says its AI Power User Index combines frequency, breadth, and depth of engagement and relates engagement to performance ratings before and after the tool arrived.
- The article says managers were three times more likely to be power users and teams whose managers were power users were twice as likely to be, highlighting plausible role and team selection differences that need analysis.
Preserve the claim's exact estimand
A buyer should first obtain the analysis population, enrollment and observation dates, eligible and excluded employees, organization levels and functions, baseline and follow-up review cycles, performance-band definitions, transition counts, missing outcomes, and the model used to calculate the 31% higher odds. Ask for absolute probabilities by group and uncertainty intervals; an odds ratio cannot be read as a probability-point change, and a move between company-specific bands may not translate into another employer's outcome. Confirm whether the reported comparison is power users versus all others, another engagement band, or a modeled contrast, and whether people could move down, remain, leave, or lack a rating. Preserve the provider, employer, and analyst roles and any publication or commercial interest.
The accountable team should translate this point into a named workflow, affected population, source data, human owner, approval right, exception path, retained evidence, and review date. That translation is what separates an interesting AI development from a decision that can be governed and evaluated.
Audit the exposure before interpreting it
The AI Power User Index combines frequency, breadth, and depth, so the buyer needs definitions, feature construction, thresholds, weights, time window, validation, sensitivity, and the distribution of each component. Depth based on how much of one's thinking appears in an exchange may encode writing style, job type, language, seniority, workload, access, or comfort disclosing context rather than coaching dose alone. Determine whether exposure was calculated before the outcome, whether review-related sessions were included, how inactive or partial users were handled, and whether people could become power users because they were already motivated, supported, or performing well. Inspect privacy-preserving aggregate evidence; do not request individual coaching content merely to validate engagement. The metric must be reproducible without turning confidential conversations into performance surveillance.
The accountable team should translate this point into a named workflow, affected population, source data, human owner, approval right, exception path, retained evidence, and review date. That translation is what separates an interesting AI development from a decision that can be governed and evaluated.
Test selection and plausible confounding
Power use was not described as randomly assigned. Ask how the study addressed prior performance and trajectory, manager status, team, function, level, tenure, location, workload, development opportunity, access timing, manager encouragement, concurrent talent programs, organizational change, rating calibration, attrition, and missing data. The provider's finding that managers and teams with power-user managers were more likely to be power users makes managerial and team context especially important. Request adjusted and unadjusted estimates, balance diagnostics, sensitivity analyses, alternative exposure thresholds, negative controls where appropriate, and results across material subgroups. Even a stable adjusted association may reflect unmeasured motivation or opportunity. Use causal language only if the design and assumptions support it; otherwise describe the observed relationship.
The accountable team should translate this point into a named workflow, affected population, source data, human owner, approval right, exception path, retained evidence, and review date. That translation is what separates an interesting AI development from a decision that can be governed and evaluated.
Design prospective buyer evidence
For a buyer pilot, pre-register the coaching purpose, eligible population, assignment or rollout method, baseline, primary and harm outcomes, exposure measure, privacy boundary, analysis, subgroup protections, and stop criteria before results are visible. Separate adoption from coaching quality and from employment evaluation; managers and decision-makers should not gain access to private session content. Measure access, engagement, self-efficacy, behavior evidence, manager and peer observations where appropriate, work outcomes, rating-process changes, adverse experiences, differential participation, attrition, and total cost. A stepped or randomized design may be possible, but only with appropriate workforce and ethics review. Valence's study is a relevant provider signal that moves the question beyond logins. It is not a substitute for method disclosure, independent scrutiny, or evidence that the buyer's configured system caused a fair and useful outcome.
The accountable team should translate this point into a named workflow, affected population, source data, human owner, approval right, exception path, retained evidence, and review date. That translation is what separates an interesting AI development from a decision that can be governed and evaluated.
Limitations and unknowns
Valence is the provider and study source. Its August 18, 2026 article reports a study of 11,000 employees at one large technology company, an AI Power User Index based on frequency, breadth, and depth, and 31% higher odds of moving up a performance band among power users. It does not publish in the article the complete protocol, eligibility and exclusions, group sizes, absolute transition probabilities, uncertainty, model specification, index thresholds and validation, missingness and attrition, adjustment set, balance and sensitivity analyses, subgroup estimates, independent replication, or evidence that engagement caused performance movement. Current full methods and aggregate results, privacy and employment-use boundaries, independent statistical review, prospective buyer design, participant evidence, and qualified coaching, people analytics, statistics, employment, accessibility, privacy, security, procurement, ethics, and legal review control.
Decision test
Ask whether the source changes the decision itself, the evidence required, the implementation sequence, or only the language used to describe an existing capability. Record which claims are directly supported, which are provider statements, which require an independent test, and which remain unknown. A source-linked review should make uncertainty easier to see, not bury it inside a blended score.
Questions to take into review
The publication supports research and executive decision preparation. It does not provide legal, financial, accounting, employment, clinical, cybersecurity, investment, procurement, or implementation advice.