Site Logo
All articles
Commentary on News

Process-Focused Assessment: Assessing Coding with AI

Evidano5 min read

Teaching teams face a new problem: finished code no longer proves learning because generative AI can produce polished artifacts. The primary keyword for this post is "assessing coding with AI, " aimed at instructors and education researchers who must verify understanding rather than artifacts. This post distills Eric Freeman's July 28, 2026 O'Reilly Radar analysis and gives concrete, AI-enabled qualitative research workflows that instructors can use to evaluate learning and scale feedback.

Key Takeaways

According to O'Reilly Radar (July 28, 2026), Eric Freeman argues that assessment must shift from finished artifacts to visible process: show drafts, make in-class work public, and use AI as a conversational assessor.

  • Freeman reported on July 28, 2026 that finished programs often reflect prompts more than student thinking, creating an assessment gap.
  • A Stanford HAI study (July 10, 2023) found that 61.22% of TOEFL essays were falsely flagged as AI-generated when detectors were applied, showing detector unreliability.
  • Freeman documents public, performative practices such as an algorave on November 20, 2025 that force students to demonstrate live fluency rather than submit polished artifacts.

What happened and how the studio model works

Eric Freeman reported in O'Reilly Radar on July 28, 2026 that generative AI has made finished student code a weaker signal of learning because the code can reflect prompts rather than reasoning.

Freeman described three concrete classroom responses used at the University of Texas at Austin's Arts and Entertainment Technologies Department: public studio work, role-inverted AI conversational assessment, and live coding performances.

"A finished program now tells us more about a student’s prompts than their ideas, " Eric Freeman wrote in O'Reilly Radar (July 28, 2026).

Freeman notes that the studio model collects process evidence: sketches, revisions, prompt histories, and oral explanations, which together provide a richer qualitative record than a final submission alone.

Findings Snapshot

DateMetric / EventValue / DetailImplication
July 28, 2026O'Reilly Radar articleEric Freeman: Teaching Coding When AI Can Write the CodeArgues assessment must focus on process and explanation
July 10, 2023Stanford HAI study61.22% of TOEFL essays flagged as AI-generatedShows AI detectors produce high false positives for nonnative writers
2023 (January–July)OpenAI classifier statusOpenAI retired its AI Text Classifier in 2023Illustrates limits of detector-first assessment strategies
November 20, 2025Algorave performanceAudioPixel Collider live coding eventLive performance forces on-the-spot fluency and visible process

Implications for educators and assessment designers

The direct answer is that assessment must measure process, explanation, revision, and recovery, not just final artifacts, as argued by Eric Freeman in O'Reilly Radar (July 28, 2026).

  • Design in-class, public milestones so students must explain choices and show drafts, which Freeman implemented at UT Austin starting in 2024 and reported in 2026.
  • Use AI as an active conversational assessor that produces transcripts and rubric-aligned summaries, as Freeman described with the Vera Molnár avatar chatbot piloted in 2025.
  • Introduce performative assessments such as live coding or demonstrations, exemplified by the AudioPixel Collider algorave on November 20, 2025, to reveal fluency under pressure.

Freeman cautions that these are experimental approaches and not universal solutions; he notes that his observations are based on small-scale experiments rather than controlled trials.

How Evidano helps when assessing coding with AI

Problem: Finished artifacts hide process → Solution: Collect process evidence

Evidano is an AI-powered qualitative data analysis platform that helps researchers analyze interviews, open-ended surveys, and documents.

Evidano can ingest prompt histories, code revision logs, classroom transcripts, and AI-chat transcripts, then generate thematic and frequency analyses that surface where students struggled and how their reasoning evolved. See Evidano features for relevant tools.

Problem: TAs need scalable, rubric-aligned feedback → Solution: AI-assisted rubric coding

Freeman reported that an AI rubric reviewer helped teaching assistants evaluate object-oriented projects in 2025, producing summaries that TAs used to make final grading decisions (O'Reilly Radar, July 28, 2026).

Evidano can apply rubric-driven automated coding to transcripts and submissions, provide cross-segment comparisons, and produce visualizations so TAs focus on judgment rather than low-level verification. See Evidano AI chatbot use cases for conversational assessment support.

Problem: Need verifiable conversational records → Solution: Secure transcription and analysis

Freeman described using AI avatars to converse with students and produce assessment transcripts (O'Reilly Radar, July 28, 2026).

Evidano provides transcription with custom dictionaries and PII redaction plus encrypted storage, enabling instructors to analyze conversation transcripts and extract rubric matches, common misconceptions, and intervention points. Learn about our speech-to-text capabilities.

FAQ: assessing coding with AI

How can instructors verify student understanding when AI can write code?

Answer: Instructors should verify understanding by assessing process evidence, not just final code.

Supporting detail: Freeman recommended studio-style public work, conversational AI assessments, and live performance as ways to collect observable reasoning and revisions (O'Reilly Radar, July 28, 2026).

Are AI detectors reliable for grading or honor-code enforcement?

Answer: No, AI detectors are not reliable enough to serve as the primary enforcement tool.

Supporting detail: The Stanford HAI study (July 10, 2023) found 61.22% false positives in TOEFL essays, and OpenAI retired its classifier in 2023, which Freeman cites as evidence against detector-first policies (O'Reilly Radar, July 28, 2026).

Can live coding and public studio practices scale for large classes?

Answer: They can scale with AI-enabled transcription and rubric automation to reduce reviewer workload.

Supporting detail: Freeman described small experiments that paired AI rubric summaries with TA review to scale feedback; platforms that capture transcripts and automate thematic coding make studio evidence review practical at larger scale (O'Reilly Radar, July 28, 2026).

How can qualitative analysis tools support rubric-aligned grading?

Answer: Qualitative tools can map transcripts and revision histories to rubric categories and surface evidence for each criterion.

Supporting detail: Freeman described AI summaries aligned to rubrics that TAs used for grading in 2025 experiments, and Evidano can produce the same structured outputs from chat transcripts and code change logs.

Conclusion & Next Steps

The bottom line from Eric Freeman in O'Reilly Radar (July 28, 2026) is clear: as AI makes polished code easier to produce, assessment must focus on process, explanation, revision, and performance.

Instructors can combine studio practices, AI conversational assessments, and live demonstrations to collect verifiable evidence of understanding.

To pilot these methods, capture transcripts, prompt histories, and revision logs, then run thematic and rubric analyses to surface learning gaps.

Get started by exploring how an AI-native qualitative platform can ingest transcripts and code logs and produce rubric-aligned summaries: Try Evidano for free.

Company
About
Newsletter

Product updates, research, and tips — straight to your inbox.

© Evidano, All Rights Reserved.