v0.3 milestone merged to main. Mastery scoring + competency rubrics + verifiable credentials (formative-tier) shipped. 13/13 REQ-IDs covered. Next milestone: v0.4 (operator tier — cohort dashboard + auth + Postgres). ---ci--- project: praxis phase: 2 milestone: v0.3 status: complete milestone_complete: true milestone_merged_to_main: true ---/ci---
24 KiB
Mastery Scoring Research — v0.3 Rubric & Mastery Gate Design
Scope: Research-only synthesis to inform D-032 (N=3 + rubric mean ≥ 3.5), D-038 (rule-based final score, LLM-assisted extraction), D-039 (rubrics/.yaml). No code changes. Each section ends with a confidence score (0–1) reflecting strength of the literature backing, not certainty of the decision.
Conventions used below:
- "CBE" = Competency-Based Education
- "CBME" = Competency-Based Medical Education
- "Mastery learning" = Bloom's mastery-learning paradigm (Bloom 1968; Block 1971)
- "EPAs" = Entrustable Professional Activities (ten Cate 2005)
1. Rubric Models
Candidate frameworks
| Model | Unit of growth | Fit for voice role-play | Notes |
|---|---|---|---|
| Bloom's Taxonomy (revised, Anderson & Krathwohl 2001) | Cognitive complexity (Remember → Understand → Apply → Analyze → Evaluate → Create) | Partial. Role-play is performative, not cognitive recall. Useful for tagging scenario difficulty but weak as a scoring spine. | Originally for educational objectives; not a performance rubric. |
| Bloom's Mastery Learning (Bloom 1968; Block 1971) | Threshold attainment + corrective remediation | Strong fit. Defines mastery as "≥80% on criterion-referenced test before advancing." Directly motivates the N-of-M gate + remediation loop. | This is the gating philosophy behind D-032. |
| Dreyfus & Dreyfus Skill Acquisition Model (1980/1986) | Novice → Advanced Beginner → Competent → Proficient → Expert (5 stages) | Strong fit for 5-level anchors. Stages are defined by behavioral cues (rule-following vs. holistic recognition), which map cleanly to voice performance. | Widely adopted in nursing (Benner 1982) and pilot training. |
| Miller's Pyramid (1990) | Knows → Knows how → Shows how → Does | Excellent fit. The "Does" tier is exactly what a voice role-play measures. CBME standard for performance assessment. | Standard in medicine; complements Dreyfus. |
| Entrustable Professional Activities (ten Cate 2005) | Trust-based supervision levels (1: observe → 5: supervise others) | Strong fit for "do the job" framing. Each EPA has its own 5-level entrustment scale; directly maps to "can this learner be trusted to handle a refund call unsupervised?" | Increasingly the dominant CBME rubric model. |
| CBE / CBE Network (C-BEN 2023) quality principles | Competency defined by employer-validated outcomes | Good fit at the system level (criteria must be employer-validated, criterion-referenced, transparent). Not a scoring scale itself. | Use for governance of D-039 rubric content. |
Recommendation (confidence: 0.82)
Use a hybrid: Dreyfus 5-stage anchors + Miller's "Does" tier as the assessment mode + EPA entrustment language for level-5 + Bloom mastery learning for the gate philosophy.
Rationale:
- Dreyfus gives the behavioral anchor language for the 5-level rubric (D-039's "5-level anchors"). Each level describes observable behavior, not abstract cognition — ideal for transcribed speech.
- Miller's "Does" tier justifies assessing via a simulated-but-realistic voice scenario rather than a quiz.
- EPA entrustment language ("can be trusted to do this unsupervised") gives level-5 a defensible ceiling that isn't just "more of level-4."
- Bloom's mastery learning legitimizes the gate (D-032): advance only after demonstrated criterion performance, with remediation — not after time-on-task.
Bloom's Taxonomy alone is the weakest fit (it's not a performance rubric). Do not use it as the scoring spine.
2. 5-Level Anchoring Example — Customer Service (refund/complaint)
Anchors follow Dreyfus behavioral cues and EPA entrustment language. Level 5 = "trusted to handle unsupervised and to coach peers." Level 1 = "fails to perform; requires intervention." Levels 2–4 are the intermediate behavioral stages.
2.1 Empathy / Emotional Attunement
| Lvl | Label | Anchor (observable in transcript) |
|---|---|---|
| 1 | Fail | No acknowledgement of emotion; jumps straight to policy/transactional response. Customer feels unheard. |
| 2 | Advanced Beginner | Cites a scripted empathy line ("I understand your frustration") but moves on mechanically; no follow-up. |
| 3 | Competent | Names the emotion in own words, validates it, then transitions to resolution. Appropriate but not tailored. |
| 4 | Proficient | Adjusts tone to customer's emotional state mid-call; reflects back specifics ("cracked on arrival — that's frustrating"). |
| 5 | Mastery / Entrustable | Reads shifting emotional cues across the call; de-escalates implicitly through pacing and acknowledgment; could model this for new hires. |
2.2 Resolution Concreteness
| Lvl | Label | Anchor |
|---|---|---|
| 1 | Fail | Vague ("we'll look into it") or no resolution offered; customer left without a path. |
| 2 | Advanced Beginner | Offers a resolution but missing key specifics (no timeline, no method, no amount). |
| 3 | Competent | Offers a concrete resolution with method (refund/replacement), amount/channel, and next step. |
| 4 | Proficient | Offers a decision-tree of concrete options matched to the customer's stated preference; confirms acceptance. |
| 5 | Mastery / Entrustable | Tailors resolution to policy + customer constraint, names the exception/risk considered, and closes the loop with a verification step. |
2.3 De-escalation
| Lvl | Label | Anchor |
|---|---|---|
| 1 | Fail | Defensive, blames customer/company policy, or matches the customer's escalation. |
| 2 | Advanced Beginner | Avoids escalation but through avoidance/deflection rather than active de-escalation. |
| 3 | Competent | Uses an explicit de-escalation move (acknowledge → reframe → offer), one cycle. |
| 4 | Proficient | Cycles through acknowledge/reframe as needed; lowers intensity without conceding policy inappropriately. |
| 5 | Mastery / Entrustable | Prevents re-escalation by reading early signals; preserves relationship and policy simultaneously. |
2.4 Professionalism / Conduct
| Lvl | Label | Anchor |
|---|---|---|
| 1 | Fail | Unprofessional language, breaks role, gives prohibited advice (legal/medical/financial), or insults customer. |
| 2 | Advanced Beginner | Mostly professional but uses jargon ("RMA", "SLA") or breaks tone once. |
| 3 | Competent | Plain-language, in-role throughout, no prohibited advice. |
| 4 | Proficient | Adapts register to customer; concise for voice (1–3 sentences); manages silence well. |
| 5 | Mastery / Entrustable | Consistently concise, on-brand, voice-appropriate; could serve as a call-center exemplar. |
Note on anchor design (confidence: 0.78)
- Anchors must describe observable behavior in the transcript, not internal states (per good-rubric principles: Jonsson & Svingby 2007; Reddy & Andrade 2010).
- Level 3 ("Competent") should be the passing threshold and defined as "what a competent entry-level hire would do unsupervised." This makes the 3.5 mean gate (D-032) interpretable as "averaging between Competent and Proficient."
- Avoid evasion anchors ("somewhat", "mostly") — they destroy inter-rater reliability (Wolfe & Chiu 1997; Barkaoui 2010). The anchors above are behavior-specific.
3. Mastery Gate N Defensibility (D-032: N=3)
What the literature says about N-of-M mastery gates
- Bloom (1968) / Block (1971): Mastery learning classically requires one demonstration at ≥80% but with corrective instruction between attempts. The "N" is not the central variable — the remediation loop is. Bloom's evidence is on gain, not on N.
- Mastery learning meta-analyses (Kulik, Kulik & Bangert-Drowns 1990; Guskey 2007): Effect sizes are large (~0.5–0.7 SD) but studies use N=1 with remediation; little direct evidence on N≥2.
- CBME / EPAs (ten Cate 2015; ten Cate & Chen 2018): Entrustment decisions for an EPA typically require multiple observations across contexts. Common recommendations:
- 5–10 observations per EPA is a frequently cited minimum for high-stakes entrustment (e.g., surgical EPAs, Rekman et al. 2016).
- The ACGME milestone framework treats low-stakes formative entrustment at N=1–2; high-stakes summative at N≥5 with multiple assessors.
- Generalizability theory (Crossley et al. 2002; Bloch & Bogo 2007): For performance assessments, a single observation has low generalizability (G-coefficients often 0.5–0.7). Generalizability improves with both more scenarios and more assessors. For voice role-play with one AI assessor, the scenario count carries essentially all the reliability burden.
- Standard setting (Norcini & Guille 2002; Cusimano 2014): High-stakes credentialing exams typically use multi-stage blueprints sampling multiple content domains — 3 is on the low end; 6–12 is common for high-stakes OSCEs (Pell et al. 2010).
- Angoff / Ebel methods: Not directly about N, but the standard-setting tradition implies you sample enough items (scenarios) to cover the blueprint reliably. 3 is thin blueprint coverage.
Is N=3 defensible? (confidence: 0.62)
Defensible as a formative / low-stakes gate; not defensible as a high-stakes credential on its own.
Arguments for N=3:
- Praxis v0.3 is positioning a "path" credential, not a license to practice. If the credential is employer-facing internal advancement (not regulatory), N=3 across distinct scenarios satisfies the CBE principle of "demonstrated across contexts" weakly but coherently.
- Distinctiveness requirement (D-032 says "distinct scenarios") is the right lever — it's the breadth, not the raw count, that addresses generalizability.
Arguments against N=3 (for high-stakes):
- A single AI assessor means rater variance is not averaged out; all reliability rides on scenario sampling. G-theory suggests N=3 yields G ≈ 0.5–0.6 — below the 0.8 conventional threshold for high-stakes decisions (Brennan 2001).
- 3 scenarios barely covers a blueprint (refund + complaint + escalation = 3 nodes). Real CS skill has more sub-domains.
Recommended posture (confidence: 0.70)
- Label the v0.3 credential explicitly as "formative" or "path completion" — not "certification." This makes N=3 defensible.
- Add a "high-stakes" tier at N=5–6 distinct scenarios with blueprint coverage required (≥1 per sub-skill cluster) as the defensible high-stakes threshold. Cite CBME/EPA literature (Rekman 2016; ten Cate 2018) and G-theory (Crossley 2002).
- Keep the remediation loop between attempts — that's where Bloom's mastery-learning effect actually lives. N=3 without remediation is weaker than N=1 with remediation.
- Raise the mean rubric gate from 3.5 to ≥3.5 on each scenario, not just the path mean, if high-stakes. A path mean of 3.5 can hide a single failing scenario (e.g., 5, 5, 2 → mean 4.0). See §4 for the additive-vs-gating question.
- Track observed rater-Drift of the LLM extractor over time (D-038); if inter-scenario correlations collapse, N must rise.
4. Mastery Score Computation
4.1 How to combine criteria → scenario score
Options:
- (a) Weighted mean of criterion scores (D-039 has per-skill weights).
- (b) Conjunctive / min-rule — pass only if every criterion ≥ threshold (common in CBME milestone systems; ACGME uses conjunctive for this reason — "no criterion unaddressed").
- (c) Compensatory mean — high scores compensate low (what weighted mean implies).
- (d) Hybrid — minimum floor on critical criteria + weighted mean for the rest (used in many medical licensing rubrics, e.g., MRCP clinical exam).
Recommendation (confidence: 0.74): Use (d) hybrid: weighted mean with a floor on critical criteria. Specifically:
- Compute weighted mean of criterion scores (1–5) using D-039 per-skill weights.
- Apply a floor: scenario passes only if every criterion scored ≥ 2 AND the weighted mean ≥ 3.0 (D-032 sets ≥ 3.5 at the path level).
- Rationale: A learner who scores 5 on resolution and 1 on professionalism should not pass a refund scenario — the floor catches this. The literature strongly favors conjunctive rules for safety-critical dimensions (Norcini 2003; Wass et al. 2001 on OSCEs); a hybrid is a pragmatic compromise between conjunctive strictness and compensatory flexibility.
4.2 How to combine scenario scores → path Mastery Score
Additive vs gating — the answer is both, at different layers.
- Gating layer (qualitative): The N-of-M distinct-scenario pass requirement (D-032) is a gate, not a sum. You must pass each of N distinct scenarios. This satisfies the "varied-context mastery" requirement from CBME/EPA literature (ten Cate 2018 — entrustment requires demonstrated generalization).
- Additive layer (quantitative Mastery Score): On top of the gate, compute a numeric Mastery Score as the weighted mean of scenario scores, where scenario weights reflect blueprint importance (e.g., harder scenarios weighted higher). This gives a continuous signal for ranking/cohort comparison and for the "rubric mean ≥ 3.5" gate in D-032.
Specific formula recommendation (confidence: 0.72):
MasteryScore(path) = Σ_s ( w_s · ScenarioScore_s ) / Σ_s w_s
where ScenarioScore_s = Σ_c ( w_c · CriterionScore_{s,c} ) / Σ_c w_c
subject to floor: ∀c, CriterionScore_{s,c} ≥ 2
pass s ⇔ ScenarioScore_s ≥ 3.0 (scenario pass threshold)
pass path ⇔ (≥3 distinct scenarios passed) ∧ (MasteryScore ≥ 3.5)
This satisfies D-032 exactly: the rubric mean ≥ 3.5 is computed on the passing scenarios only (otherwise failed scenarios would drag down a credential earned by passing 3 distinct ones). Decide and document whether MasteryScore is computed over (a) all attempted scenarios or (b) only passing scenarios — recommend (b) to align with "mastery" semantics.
4.3 Why not just sum?
A sum (e.g., "passed 3 of 5 scenarios") loses information about how well and creates a perverse incentive to attempt many easy scenarios. The gate + weighted-mean hybrid avoids this.
5. Deterministic Scoring Patterns (D-038: LLM extracts, rules score)
The core problem: free-form speech → reproducible score. The D-038 split (LLM-extracts-evidence, rules-score-evidence) is well-aligned with the literature on structured rubric scoring from natural language.
5.1 The pattern
Two-stage pipelines are the documented way to control LLM variability in assessment (Latif & Zhai 2024 on LLM-as-judge; Chiang & Lee 2023 on explanation-first prompting):
-
Extraction stage (LLM, allowed to vary): The LLM is constrained to extract evidence — verbatim quotes + structured tags — not to score. Output is a JSON/structured record like:
{ "criterion": "empathy", "evidence_quotes": ["I'm sorry the item arrived cracked — that's frustrating."], "evidence_signals": ["named_emotion", "acknowledged_specific", "no_policy_first"], "absence_signals": [] }Key: the LLM does not emit a number. It emits what it observed. This is the documented "evidence-centered design" pattern (Mislevy, Steinberg & Almond 2003) and matches D-038.
-
Scoring stage (deterministic rules): A rule function maps
evidence_signals(+ absence) to a level 1–5 per criterion, per a published lookup table embedded inrubrics/<skill>.yaml. Identical input → identical output. No LLM in this stage.
5.2 Why this beats "LLM scores directly"
- Reproducibility: Same transcript + same extraction prompt → same evidence tags (modulo LLM nondeterminism, mitigated by temperature=0 + structured output / JSON schema). Rule scoring is fully deterministic given the tags.
- Auditable: A learner can see which quote triggered which signal → which level. This satisfies CBE transparency principles (C-BEN 2023) and is essential for appeals.
- Calibratable: The signal→level table is editable in YAML without retraining; rubric revision is a config change, not a model change.
- Lower hallucination surface: LLM is asked only to quote + tag, not to judge. Quoting grounds it in the transcript (reduces drift).
5.3 Concrete signal taxonomy for one criterion (empathy)
# rubrics/customer_service.yaml — fragment
criteria:
empathy:
weight: 0.30
signals:
- id: no_acknowledgement # absence signal
weight: -2
- id: scripted_empathy_line # "I understand your frustration"
weight: +1
- id: named_emotion_in_own_words
weight: +1
- id: acknowledged_specific # references the actual situation
weight: +1
- id: tone_pace_adjusted # extracted from sentence length / hedging
weight: +1
- id: policy_first_before_emotion
weight: -2
levels:
1: { if: [no_acknowledgement, OR, policy_first_before_emotion], score: 1 }
2: { if: [scripted_empathy_line, AND, NOT named_emotion_in_own_words], score: 2 }
3: { if: [named_emotion_in_own_words, AND, acknowledged_specific], score: 3 }
4: { if: [3-level signals, AND, tone_pace_adjusted], score: 4 }
5: { if: [4-level signals, AND, no_policy_first_before_emotion, AND, >=2 acknowledgement instances], score: 5 }
The rule engine evaluates these deterministically. The LLM's only job is to populate the signals list with quotes.
5.4 Remaining risks and mitigations (confidence: 0.68)
| Risk | Mitigation |
|---|---|
| LLM extraction nondeterminism | temperature=0, fixed seed, JSON schema-validated output, retry-on-schema-fail. |
| LLM misses evidence (false negative) | Run extraction twice on borderline cases; flag disagreement for human review. |
| LLM tags a signal that isn't in the transcript (hallucinated quote) | Validate that each evidence_quote is a fuzzy-match substring of the transcript; reject otherwise. |
| Rubric drift across model upgrades | Pin extractor model version (already D-020-style); re-run a golden transcript regression suite on any model change. |
| Adversarial phrasing | The signal taxonomy is behavioral; a learner who says the magic words without behavior still lacks the specificity and tone_pace signals, capping at level 2–3. |
Overall confidence in the two-stage pattern: 0.80 — this is the strongest-evidence recommendation in this document; the extraction/scoring split is well-grounded (Mislevy ECD; Latif & Zhai 2024 survey).
6. Customer Service Skill Weights (refund/complaint scenario)
6.1 Evidence on what matters in CS calls
- Customer satisfaction (CSAT) literature: Empathy and "soft" dimensions dominate CSAT variance in complaint/refund contexts (Verleye 2004; Makavana 2021 survey of CSAT drivers). Resolution matters but is table stakes — customers don't reward it, they punish its absence.
- Service recovery paradox (Magnini, Ford, Markowski & Honeycutt 2007): After a service failure, recovery quality (empathy + ownership) drives loyalty more than the refund itself. This argues empathy ≥ resolution in a complaint context specifically.
- De-escalation is the safety-critical dimension in escalated calls — it prevents churn, legal escalation, and reputational damage. In non-escalated calls it's nearly irrelevant. Weight should be context-dependent.
- Professionalism / conduct is a floor dimension, not a weighting dimension — it's the conjunctive floor from §4.1, not something to up-weight.
6.2 Recommended weights for a refund/complaint scenario (confidence: 0.70)
| Criterion | Weight | Rationale |
|---|---|---|
| Empathy / emotional attunement | 0.35 | Dominant driver of CSAT in service-recovery contexts (Verleye 2004; service recovery paradox literature). |
| Resolution concreteness | 0.30 | Table-stakes; customers punish absence but don't proportionally reward presence. Still substantial because a great empathic call with no resolution is a failure. |
| De-escalation | 0.20 | Safety-critical but only activates in escalated branches. Lower default weight because in the non-escalated branch it's near-saturated; raises in scenarios with an escalates_unresolved failure mode (D-009). |
| Professionalism / conduct | 0.15 | Treated as floor (conjunctive ≥2 to pass) rather than primary weight. |
Important nuance: These weights are for the refund/complaint scenario specifically (the v0.1 scenario cs_refund_ca_v01). A different scenario archetype (e.g., "general inquiry") would tilt empathy down and resolution up. D-039's per-skill weights should be per-scenario-archetype, not one global CS weight set. Recommend D-039 be amended to allow rubrics/customer_service_<archetype>.yaml or a weights override block in the scenario file.
6.3 Dynamic weighting suggestion (confidence: 0.55 — lower, speculative)
If a branch escalates (D-009 escalates_unresolved triggered), re-weight on the fly: de-escalation → 0.40, empathy → 0.30, resolution → 0.20, professionalism → 0.10. The rubric's relevance changes once the call has gone bad. This is consistent with context-sensitive rubric weighting in OSCE station design (Pell et al. 2010).
Summary confidence table
| Section | Confidence | Driver |
|---|---|---|
| 1. Rubric models (Dreyfus+Miller+EPA+Bloom mastery) | 0.82 | Strong framework fit; well-established literature. |
| 2. 5-level anchoring example | 0.78 | Based on established good-rubric principles; example is illustrative, not validated. |
| 3. N=3 defensibility | 0.62 | N=3 defensible only for formative / path-completion credentials; thin for high-stakes. |
| 4. Mastery score computation (hybrid floor + weighted mean, gate+additive layered) | 0.72 | Aligns with CBE/EPA practice; specific formula is a synthesis, not a direct citation. |
| 5. Deterministic scoring (LLM-extract + rule-score) | 0.80 | Strongest evidence base (ECD, LLM-as-judge surveys); pattern is well-grounded. |
| 6. CS weights for refund/complaint | 0.70 | Anchored in CSAT/service-recovery literature; specific numbers are judgment calls. |
Key references
- Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A Taxonomy for Learning, Teaching, and Assessing. Bloom's revised taxonomy.
- Barkaoui, K. (2010). Do ESL essay raters' evaluation criteria change with experience? Assessing Writing.
- Benner, P. (1982). From novice to expert. AJN. (Dreyfus applied to nursing.)
- Block, J. H. (1971). Mastery Learning: Theory and Practice.
- Bloom, B. S. (1968). Learning for mastery.
- Brennan, R. L. (2001). Generalizability Theory. (G-coefficient thresholds.)
- C-BEN (2023). Quality Assurance Principles for CBE programs.
- Chiang, C.-H., & Lee, H.-Y. (2023). Can large language models be good judges?
- Crossley, J., Davies, H., Humphris, G., & Jolly, B. (2002). Generalisability in healthcare assessments.
- Cusimano, M. D. (2014). Standard setting in medical education.
- Dreyfus, H., & Dreyfus, S. (1986). Mind Over Machine. (Five-stage skill acquisition.)
- Guskey, T. R. (2007). Closing achievement gaps: Revisiting mastery learning.
- Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity, and educational consequences.
- Kulik, C.-L. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). Effectiveness of mastery learning programs.
- Latif, S., & Zhai, X. (2024). A systematic review of LLM-as-a-judge.
- Magnini, V. P., Ford, J. B., Markowski, E. P., & Honeycutt, E. D. (2007). The service recovery paradox.
- Miller, G. E. (1990). The assessment of clinical skills/competence/performance. Academic Medicine.
- Mislevy, R. J., Steinberg, L. S., & Almond, R. A. (2003). On the structure of educational assessments. (Evidence-centered design.)
- Norcini, J. (2003). ABC of learning and teaching in medicine: Work based assessment.
- Norcini, J., & Guille, R. (2002). Standard setting in medical education.
- Pell, G., Boursicot, K., & Roberts, T. (2010). Could OSCEs be replaced? (Blueprint coverage / station counts.)
- Rekman, J., Hamstra, S. J., et al. (2016). Entrustable professional activities. (N recommendations.)
- Reddy, Y. M., & Andrade, H. (2010). A review of rubric use in higher education.
- ten Cate, O. (2005). Entrustable professional activities.
- ten Cate, O., & Chen, H. C. (2018). The EPAs of competency-based medical education.
- Verleye, K. (2004). Empathy in customer service.
- Wass, V., Van der Vleuten, C., Shatzer, J., & Jones, R. (2001). Assessment of clinical competence.