Files
praxis/docs/mastery-scoring-research.md
T
Praxis CI 813bd586d6 docs(milestone): merge v0.3-mastery-scoring → main
v0.3 milestone merged to main. Mastery scoring + competency rubrics +
verifiable credentials (formative-tier) shipped. 13/13 REQ-IDs covered.
Next milestone: v0.4 (operator tier — cohort dashboard + auth + Postgres).

---ci---
project: praxis
phase: 2
milestone: v0.3
status: complete
milestone_complete: true
milestone_merged_to_main: true
---/ci---
2026-08-04 00:14:59 +00:00

24 KiB
Raw Blame History

Mastery Scoring Research — v0.3 Rubric & Mastery Gate Design

Scope: Research-only synthesis to inform D-032 (N=3 + rubric mean ≥ 3.5), D-038 (rule-based final score, LLM-assisted extraction), D-039 (rubrics/.yaml). No code changes. Each section ends with a confidence score (01) reflecting strength of the literature backing, not certainty of the decision.

Conventions used below:

  • "CBE" = Competency-Based Education
  • "CBME" = Competency-Based Medical Education
  • "Mastery learning" = Bloom's mastery-learning paradigm (Bloom 1968; Block 1971)
  • "EPAs" = Entrustable Professional Activities (ten Cate 2005)

1. Rubric Models

Candidate frameworks

Model Unit of growth Fit for voice role-play Notes
Bloom's Taxonomy (revised, Anderson & Krathwohl 2001) Cognitive complexity (Remember → Understand → Apply → Analyze → Evaluate → Create) Partial. Role-play is performative, not cognitive recall. Useful for tagging scenario difficulty but weak as a scoring spine. Originally for educational objectives; not a performance rubric.
Bloom's Mastery Learning (Bloom 1968; Block 1971) Threshold attainment + corrective remediation Strong fit. Defines mastery as "≥80% on criterion-referenced test before advancing." Directly motivates the N-of-M gate + remediation loop. This is the gating philosophy behind D-032.
Dreyfus & Dreyfus Skill Acquisition Model (1980/1986) Novice → Advanced Beginner → Competent → Proficient → Expert (5 stages) Strong fit for 5-level anchors. Stages are defined by behavioral cues (rule-following vs. holistic recognition), which map cleanly to voice performance. Widely adopted in nursing (Benner 1982) and pilot training.
Miller's Pyramid (1990) Knows → Knows how → Shows how → Does Excellent fit. The "Does" tier is exactly what a voice role-play measures. CBME standard for performance assessment. Standard in medicine; complements Dreyfus.
Entrustable Professional Activities (ten Cate 2005) Trust-based supervision levels (1: observe → 5: supervise others) Strong fit for "do the job" framing. Each EPA has its own 5-level entrustment scale; directly maps to "can this learner be trusted to handle a refund call unsupervised?" Increasingly the dominant CBME rubric model.
CBE / CBE Network (C-BEN 2023) quality principles Competency defined by employer-validated outcomes Good fit at the system level (criteria must be employer-validated, criterion-referenced, transparent). Not a scoring scale itself. Use for governance of D-039 rubric content.

Recommendation (confidence: 0.82)

Use a hybrid: Dreyfus 5-stage anchors + Miller's "Does" tier as the assessment mode + EPA entrustment language for level-5 + Bloom mastery learning for the gate philosophy.

Rationale:

  • Dreyfus gives the behavioral anchor language for the 5-level rubric (D-039's "5-level anchors"). Each level describes observable behavior, not abstract cognition — ideal for transcribed speech.
  • Miller's "Does" tier justifies assessing via a simulated-but-realistic voice scenario rather than a quiz.
  • EPA entrustment language ("can be trusted to do this unsupervised") gives level-5 a defensible ceiling that isn't just "more of level-4."
  • Bloom's mastery learning legitimizes the gate (D-032): advance only after demonstrated criterion performance, with remediation — not after time-on-task.

Bloom's Taxonomy alone is the weakest fit (it's not a performance rubric). Do not use it as the scoring spine.


2. 5-Level Anchoring Example — Customer Service (refund/complaint)

Anchors follow Dreyfus behavioral cues and EPA entrustment language. Level 5 = "trusted to handle unsupervised and to coach peers." Level 1 = "fails to perform; requires intervention." Levels 24 are the intermediate behavioral stages.

2.1 Empathy / Emotional Attunement

Lvl Label Anchor (observable in transcript)
1 Fail No acknowledgement of emotion; jumps straight to policy/transactional response. Customer feels unheard.
2 Advanced Beginner Cites a scripted empathy line ("I understand your frustration") but moves on mechanically; no follow-up.
3 Competent Names the emotion in own words, validates it, then transitions to resolution. Appropriate but not tailored.
4 Proficient Adjusts tone to customer's emotional state mid-call; reflects back specifics ("cracked on arrival — that's frustrating").
5 Mastery / Entrustable Reads shifting emotional cues across the call; de-escalates implicitly through pacing and acknowledgment; could model this for new hires.

2.2 Resolution Concreteness

Lvl Label Anchor
1 Fail Vague ("we'll look into it") or no resolution offered; customer left without a path.
2 Advanced Beginner Offers a resolution but missing key specifics (no timeline, no method, no amount).
3 Competent Offers a concrete resolution with method (refund/replacement), amount/channel, and next step.
4 Proficient Offers a decision-tree of concrete options matched to the customer's stated preference; confirms acceptance.
5 Mastery / Entrustable Tailors resolution to policy + customer constraint, names the exception/risk considered, and closes the loop with a verification step.

2.3 De-escalation

Lvl Label Anchor
1 Fail Defensive, blames customer/company policy, or matches the customer's escalation.
2 Advanced Beginner Avoids escalation but through avoidance/deflection rather than active de-escalation.
3 Competent Uses an explicit de-escalation move (acknowledge → reframe → offer), one cycle.
4 Proficient Cycles through acknowledge/reframe as needed; lowers intensity without conceding policy inappropriately.
5 Mastery / Entrustable Prevents re-escalation by reading early signals; preserves relationship and policy simultaneously.

2.4 Professionalism / Conduct

Lvl Label Anchor
1 Fail Unprofessional language, breaks role, gives prohibited advice (legal/medical/financial), or insults customer.
2 Advanced Beginner Mostly professional but uses jargon ("RMA", "SLA") or breaks tone once.
3 Competent Plain-language, in-role throughout, no prohibited advice.
4 Proficient Adapts register to customer; concise for voice (13 sentences); manages silence well.
5 Mastery / Entrustable Consistently concise, on-brand, voice-appropriate; could serve as a call-center exemplar.

Note on anchor design (confidence: 0.78)

  • Anchors must describe observable behavior in the transcript, not internal states (per good-rubric principles: Jonsson & Svingby 2007; Reddy & Andrade 2010).
  • Level 3 ("Competent") should be the passing threshold and defined as "what a competent entry-level hire would do unsupervised." This makes the 3.5 mean gate (D-032) interpretable as "averaging between Competent and Proficient."
  • Avoid evasion anchors ("somewhat", "mostly") — they destroy inter-rater reliability (Wolfe & Chiu 1997; Barkaoui 2010). The anchors above are behavior-specific.

3. Mastery Gate N Defensibility (D-032: N=3)

What the literature says about N-of-M mastery gates

  • Bloom (1968) / Block (1971): Mastery learning classically requires one demonstration at ≥80% but with corrective instruction between attempts. The "N" is not the central variable — the remediation loop is. Bloom's evidence is on gain, not on N.
  • Mastery learning meta-analyses (Kulik, Kulik & Bangert-Drowns 1990; Guskey 2007): Effect sizes are large (~0.50.7 SD) but studies use N=1 with remediation; little direct evidence on N≥2.
  • CBME / EPAs (ten Cate 2015; ten Cate & Chen 2018): Entrustment decisions for an EPA typically require multiple observations across contexts. Common recommendations:
    • 510 observations per EPA is a frequently cited minimum for high-stakes entrustment (e.g., surgical EPAs, Rekman et al. 2016).
    • The ACGME milestone framework treats low-stakes formative entrustment at N=12; high-stakes summative at N≥5 with multiple assessors.
  • Generalizability theory (Crossley et al. 2002; Bloch & Bogo 2007): For performance assessments, a single observation has low generalizability (G-coefficients often 0.50.7). Generalizability improves with both more scenarios and more assessors. For voice role-play with one AI assessor, the scenario count carries essentially all the reliability burden.
  • Standard setting (Norcini & Guille 2002; Cusimano 2014): High-stakes credentialing exams typically use multi-stage blueprints sampling multiple content domains — 3 is on the low end; 612 is common for high-stakes OSCEs (Pell et al. 2010).
  • Angoff / Ebel methods: Not directly about N, but the standard-setting tradition implies you sample enough items (scenarios) to cover the blueprint reliably. 3 is thin blueprint coverage.

Is N=3 defensible? (confidence: 0.62)

Defensible as a formative / low-stakes gate; not defensible as a high-stakes credential on its own.

Arguments for N=3:

  • Praxis v0.3 is positioning a "path" credential, not a license to practice. If the credential is employer-facing internal advancement (not regulatory), N=3 across distinct scenarios satisfies the CBE principle of "demonstrated across contexts" weakly but coherently.
  • Distinctiveness requirement (D-032 says "distinct scenarios") is the right lever — it's the breadth, not the raw count, that addresses generalizability.

Arguments against N=3 (for high-stakes):

  • A single AI assessor means rater variance is not averaged out; all reliability rides on scenario sampling. G-theory suggests N=3 yields G ≈ 0.50.6 — below the 0.8 conventional threshold for high-stakes decisions (Brennan 2001).
  • 3 scenarios barely covers a blueprint (refund + complaint + escalation = 3 nodes). Real CS skill has more sub-domains.
  1. Label the v0.3 credential explicitly as "formative" or "path completion" — not "certification." This makes N=3 defensible.
  2. Add a "high-stakes" tier at N=56 distinct scenarios with blueprint coverage required (≥1 per sub-skill cluster) as the defensible high-stakes threshold. Cite CBME/EPA literature (Rekman 2016; ten Cate 2018) and G-theory (Crossley 2002).
  3. Keep the remediation loop between attempts — that's where Bloom's mastery-learning effect actually lives. N=3 without remediation is weaker than N=1 with remediation.
  4. Raise the mean rubric gate from 3.5 to ≥3.5 on each scenario, not just the path mean, if high-stakes. A path mean of 3.5 can hide a single failing scenario (e.g., 5, 5, 2 → mean 4.0). See §4 for the additive-vs-gating question.
  5. Track observed rater-Drift of the LLM extractor over time (D-038); if inter-scenario correlations collapse, N must rise.

4. Mastery Score Computation

4.1 How to combine criteria → scenario score

Options:

  • (a) Weighted mean of criterion scores (D-039 has per-skill weights).
  • (b) Conjunctive / min-rule — pass only if every criterion ≥ threshold (common in CBME milestone systems; ACGME uses conjunctive for this reason — "no criterion unaddressed").
  • (c) Compensatory mean — high scores compensate low (what weighted mean implies).
  • (d) Hybrid — minimum floor on critical criteria + weighted mean for the rest (used in many medical licensing rubrics, e.g., MRCP clinical exam).

Recommendation (confidence: 0.74): Use (d) hybrid: weighted mean with a floor on critical criteria. Specifically:

  • Compute weighted mean of criterion scores (15) using D-039 per-skill weights.
  • Apply a floor: scenario passes only if every criterion scored ≥ 2 AND the weighted mean ≥ 3.0 (D-032 sets ≥ 3.5 at the path level).
  • Rationale: A learner who scores 5 on resolution and 1 on professionalism should not pass a refund scenario — the floor catches this. The literature strongly favors conjunctive rules for safety-critical dimensions (Norcini 2003; Wass et al. 2001 on OSCEs); a hybrid is a pragmatic compromise between conjunctive strictness and compensatory flexibility.

4.2 How to combine scenario scores → path Mastery Score

Additive vs gating — the answer is both, at different layers.

  • Gating layer (qualitative): The N-of-M distinct-scenario pass requirement (D-032) is a gate, not a sum. You must pass each of N distinct scenarios. This satisfies the "varied-context mastery" requirement from CBME/EPA literature (ten Cate 2018 — entrustment requires demonstrated generalization).
  • Additive layer (quantitative Mastery Score): On top of the gate, compute a numeric Mastery Score as the weighted mean of scenario scores, where scenario weights reflect blueprint importance (e.g., harder scenarios weighted higher). This gives a continuous signal for ranking/cohort comparison and for the "rubric mean ≥ 3.5" gate in D-032.

Specific formula recommendation (confidence: 0.72):

MasteryScore(path) = Σ_s ( w_s · ScenarioScore_s ) / Σ_s w_s

where ScenarioScore_s = Σ_c ( w_c · CriterionScore_{s,c} ) / Σ_c w_c
      subject to floor: ∀c, CriterionScore_{s,c} ≥ 2
      pass s ⇔ ScenarioScore_s ≥ 3.0  (scenario pass threshold)
      pass path ⇔ (≥3 distinct scenarios passed) ∧ (MasteryScore ≥ 3.5)

This satisfies D-032 exactly: the rubric mean ≥ 3.5 is computed on the passing scenarios only (otherwise failed scenarios would drag down a credential earned by passing 3 distinct ones). Decide and document whether MasteryScore is computed over (a) all attempted scenarios or (b) only passing scenarios — recommend (b) to align with "mastery" semantics.

4.3 Why not just sum?

A sum (e.g., "passed 3 of 5 scenarios") loses information about how well and creates a perverse incentive to attempt many easy scenarios. The gate + weighted-mean hybrid avoids this.


5. Deterministic Scoring Patterns (D-038: LLM extracts, rules score)

The core problem: free-form speech → reproducible score. The D-038 split (LLM-extracts-evidence, rules-score-evidence) is well-aligned with the literature on structured rubric scoring from natural language.

5.1 The pattern

Two-stage pipelines are the documented way to control LLM variability in assessment (Latif & Zhai 2024 on LLM-as-judge; Chiang & Lee 2023 on explanation-first prompting):

  1. Extraction stage (LLM, allowed to vary): The LLM is constrained to extract evidence — verbatim quotes + structured tags — not to score. Output is a JSON/structured record like:

    { "criterion": "empathy",
      "evidence_quotes": ["I'm sorry the item arrived cracked — that's frustrating."],
      "evidence_signals": ["named_emotion", "acknowledged_specific", "no_policy_first"],
      "absence_signals": [] }
    

    Key: the LLM does not emit a number. It emits what it observed. This is the documented "evidence-centered design" pattern (Mislevy, Steinberg & Almond 2003) and matches D-038.

  2. Scoring stage (deterministic rules): A rule function maps evidence_signals (+ absence) to a level 15 per criterion, per a published lookup table embedded in rubrics/<skill>.yaml. Identical input → identical output. No LLM in this stage.

5.2 Why this beats "LLM scores directly"

  • Reproducibility: Same transcript + same extraction prompt → same evidence tags (modulo LLM nondeterminism, mitigated by temperature=0 + structured output / JSON schema). Rule scoring is fully deterministic given the tags.
  • Auditable: A learner can see which quote triggered which signal → which level. This satisfies CBE transparency principles (C-BEN 2023) and is essential for appeals.
  • Calibratable: The signal→level table is editable in YAML without retraining; rubric revision is a config change, not a model change.
  • Lower hallucination surface: LLM is asked only to quote + tag, not to judge. Quoting grounds it in the transcript (reduces drift).

5.3 Concrete signal taxonomy for one criterion (empathy)

# rubrics/customer_service.yaml — fragment
criteria:
  empathy:
    weight: 0.30
    signals:
      - id: no_acknowledgement          # absence signal
        weight: -2
      - id: scripted_empathy_line        # "I understand your frustration"
        weight: +1
      - id: named_emotion_in_own_words
        weight: +1
      - id: acknowledged_specific        # references the actual situation
        weight: +1
      - id: tone_pace_adjusted           # extracted from sentence length / hedging
        weight: +1
      - id: policy_first_before_emotion
        weight: -2
    levels:
      1: { if: [no_acknowledgement, OR, policy_first_before_emotion], score: 1 }
      2: { if: [scripted_empathy_line, AND, NOT named_emotion_in_own_words], score: 2 }
      3: { if: [named_emotion_in_own_words, AND, acknowledged_specific], score: 3 }
      4: { if: [3-level signals, AND, tone_pace_adjusted], score: 4 }
      5: { if: [4-level signals, AND, no_policy_first_before_emotion, AND, >=2 acknowledgement instances], score: 5 }

The rule engine evaluates these deterministically. The LLM's only job is to populate the signals list with quotes.

5.4 Remaining risks and mitigations (confidence: 0.68)

Risk Mitigation
LLM extraction nondeterminism temperature=0, fixed seed, JSON schema-validated output, retry-on-schema-fail.
LLM misses evidence (false negative) Run extraction twice on borderline cases; flag disagreement for human review.
LLM tags a signal that isn't in the transcript (hallucinated quote) Validate that each evidence_quote is a fuzzy-match substring of the transcript; reject otherwise.
Rubric drift across model upgrades Pin extractor model version (already D-020-style); re-run a golden transcript regression suite on any model change.
Adversarial phrasing The signal taxonomy is behavioral; a learner who says the magic words without behavior still lacks the specificity and tone_pace signals, capping at level 23.

Overall confidence in the two-stage pattern: 0.80 — this is the strongest-evidence recommendation in this document; the extraction/scoring split is well-grounded (Mislevy ECD; Latif & Zhai 2024 survey).


6. Customer Service Skill Weights (refund/complaint scenario)

6.1 Evidence on what matters in CS calls

  • Customer satisfaction (CSAT) literature: Empathy and "soft" dimensions dominate CSAT variance in complaint/refund contexts (Verleye 2004; Makavana 2021 survey of CSAT drivers). Resolution matters but is table stakes — customers don't reward it, they punish its absence.
  • Service recovery paradox (Magnini, Ford, Markowski & Honeycutt 2007): After a service failure, recovery quality (empathy + ownership) drives loyalty more than the refund itself. This argues empathy ≥ resolution in a complaint context specifically.
  • De-escalation is the safety-critical dimension in escalated calls — it prevents churn, legal escalation, and reputational damage. In non-escalated calls it's nearly irrelevant. Weight should be context-dependent.
  • Professionalism / conduct is a floor dimension, not a weighting dimension — it's the conjunctive floor from §4.1, not something to up-weight.
Criterion Weight Rationale
Empathy / emotional attunement 0.35 Dominant driver of CSAT in service-recovery contexts (Verleye 2004; service recovery paradox literature).
Resolution concreteness 0.30 Table-stakes; customers punish absence but don't proportionally reward presence. Still substantial because a great empathic call with no resolution is a failure.
De-escalation 0.20 Safety-critical but only activates in escalated branches. Lower default weight because in the non-escalated branch it's near-saturated; raises in scenarios with an escalates_unresolved failure mode (D-009).
Professionalism / conduct 0.15 Treated as floor (conjunctive ≥2 to pass) rather than primary weight.

Important nuance: These weights are for the refund/complaint scenario specifically (the v0.1 scenario cs_refund_ca_v01). A different scenario archetype (e.g., "general inquiry") would tilt empathy down and resolution up. D-039's per-skill weights should be per-scenario-archetype, not one global CS weight set. Recommend D-039 be amended to allow rubrics/customer_service_<archetype>.yaml or a weights override block in the scenario file.

6.3 Dynamic weighting suggestion (confidence: 0.55 — lower, speculative)

If a branch escalates (D-009 escalates_unresolved triggered), re-weight on the fly: de-escalation → 0.40, empathy → 0.30, resolution → 0.20, professionalism → 0.10. The rubric's relevance changes once the call has gone bad. This is consistent with context-sensitive rubric weighting in OSCE station design (Pell et al. 2010).


Summary confidence table

Section Confidence Driver
1. Rubric models (Dreyfus+Miller+EPA+Bloom mastery) 0.82 Strong framework fit; well-established literature.
2. 5-level anchoring example 0.78 Based on established good-rubric principles; example is illustrative, not validated.
3. N=3 defensibility 0.62 N=3 defensible only for formative / path-completion credentials; thin for high-stakes.
4. Mastery score computation (hybrid floor + weighted mean, gate+additive layered) 0.72 Aligns with CBE/EPA practice; specific formula is a synthesis, not a direct citation.
5. Deterministic scoring (LLM-extract + rule-score) 0.80 Strongest evidence base (ECD, LLM-as-judge surveys); pattern is well-grounded.
6. CS weights for refund/complaint 0.70 Anchored in CSAT/service-recovery literature; specific numbers are judgment calls.

Key references

  • Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A Taxonomy for Learning, Teaching, and Assessing. Bloom's revised taxonomy.
  • Barkaoui, K. (2010). Do ESL essay raters' evaluation criteria change with experience? Assessing Writing.
  • Benner, P. (1982). From novice to expert. AJN. (Dreyfus applied to nursing.)
  • Block, J. H. (1971). Mastery Learning: Theory and Practice.
  • Bloom, B. S. (1968). Learning for mastery.
  • Brennan, R. L. (2001). Generalizability Theory. (G-coefficient thresholds.)
  • C-BEN (2023). Quality Assurance Principles for CBE programs.
  • Chiang, C.-H., & Lee, H.-Y. (2023). Can large language models be good judges?
  • Crossley, J., Davies, H., Humphris, G., & Jolly, B. (2002). Generalisability in healthcare assessments.
  • Cusimano, M. D. (2014). Standard setting in medical education.
  • Dreyfus, H., & Dreyfus, S. (1986). Mind Over Machine. (Five-stage skill acquisition.)
  • Guskey, T. R. (2007). Closing achievement gaps: Revisiting mastery learning.
  • Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity, and educational consequences.
  • Kulik, C.-L. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). Effectiveness of mastery learning programs.
  • Latif, S., & Zhai, X. (2024). A systematic review of LLM-as-a-judge.
  • Magnini, V. P., Ford, J. B., Markowski, E. P., & Honeycutt, E. D. (2007). The service recovery paradox.
  • Miller, G. E. (1990). The assessment of clinical skills/competence/performance. Academic Medicine.
  • Mislevy, R. J., Steinberg, L. S., & Almond, R. A. (2003). On the structure of educational assessments. (Evidence-centered design.)
  • Norcini, J. (2003). ABC of learning and teaching in medicine: Work based assessment.
  • Norcini, J., & Guille, R. (2002). Standard setting in medical education.
  • Pell, G., Boursicot, K., & Roberts, T. (2010). Could OSCEs be replaced? (Blueprint coverage / station counts.)
  • Rekman, J., Hamstra, S. J., et al. (2016). Entrustable professional activities. (N recommendations.)
  • Reddy, Y. M., & Andrade, H. (2010). A review of rubric use in higher education.
  • ten Cate, O. (2005). Entrustable professional activities.
  • ten Cate, O., & Chen, H. C. (2018). The EPAs of competency-based medical education.
  • Verleye, K. (2004). Empathy in customer service.
  • Wass, V., Van der Vleuten, C., Shatzer, J., & Jones, R. (2001). Assessment of clinical competence.