# Mastery Scoring Research — v0.3 Rubric & Mastery Gate Design **Scope:** Research-only synthesis to inform D-032 (N=3 + rubric mean ≥ 3.5), D-038 (rule-based final score, LLM-assisted extraction), D-039 (rubrics/.yaml). No code changes. Each section ends with a confidence score (0–1) reflecting strength of the literature backing, not certainty of the decision. Conventions used below: - "CBE" = Competency-Based Education - "CBME" = Competency-Based Medical Education - "Mastery learning" = Bloom's mastery-learning paradigm (Bloom 1968; Block 1971) - "EPAs" = Entrustable Professional Activities (ten Cate 2005) --- ## 1. Rubric Models ### Candidate frameworks | Model | Unit of growth | Fit for voice role-play | Notes | |---|---|---|---| | **Bloom's Taxonomy (revised, Anderson & Krathwohl 2001)** | Cognitive complexity (Remember → Understand → Apply → Analyze → Evaluate → Create) | Partial. Role-play is *performative*, not cognitive recall. Useful for tagging scenario difficulty but weak as a scoring spine. | Originally for educational objectives; not a performance rubric. | | **Bloom's Mastery Learning (Bloom 1968; Block 1971)** | Threshold attainment + corrective remediation | Strong fit. Defines mastery as "≥80% on criterion-referenced test before advancing." Directly motivates the N-of-M gate + remediation loop. | This is the *gating* philosophy behind D-032. | | **Dreyfus & Dreyfus Skill Acquisition Model (1980/1986)** | Novice → Advanced Beginner → Competent → Proficient → Expert (5 stages) | Strong fit for 5-level anchors. Stages are defined by *behavioral cues* (rule-following vs. holistic recognition), which map cleanly to voice performance. | Widely adopted in nursing (Benner 1982) and pilot training. | | **Miller's Pyramid (1990)** | Knows → Knows how → Shows how → Does | Excellent fit. The "Does" tier is exactly what a voice role-play measures. CBME standard for performance assessment. | Standard in medicine; complements Dreyfus. | | **Entrustable Professional Activities (ten Cate 2005)** | Trust-based supervision levels (1: observe → 5: supervise others) | Strong fit for "do the job" framing. Each EPA has its own 5-level entrustment scale; directly maps to "can this learner be trusted to handle a refund call unsupervised?" | Increasingly the dominant CBME rubric model. | | **CBE / CBE Network (C-BEN 2023) quality principles** | Competency defined by employer-validated outcomes | Good fit at the *system* level (criteria must be employer-validated, criterion-referenced, transparent). Not a scoring scale itself. | Use for governance of D-039 rubric content. | ### Recommendation (confidence: **0.82**) Use a **hybrid: Dreyfus 5-stage anchors + Miller's "Does" tier as the assessment mode + EPA entrustment language for level-5 + Bloom mastery learning for the gate philosophy.** Rationale: - Dreyfus gives the *behavioral anchor language* for the 5-level rubric (D-039's "5-level anchors"). Each level describes observable behavior, not abstract cognition — ideal for transcribed speech. - Miller's "Does" tier justifies assessing via a simulated-but-realistic voice scenario rather than a quiz. - EPA entrustment language ("can be trusted to do this unsupervised") gives level-5 a defensible ceiling that isn't just "more of level-4." - Bloom's mastery learning legitimizes the **gate** (D-032): advance only after demonstrated criterion performance, with remediation — not after time-on-task. Bloom's *Taxonomy* alone is the weakest fit (it's not a performance rubric). Do not use it as the scoring spine. --- ## 2. 5-Level Anchoring Example — Customer Service (refund/complaint) Anchors follow Dreyfus behavioral cues and EPA entrustment language. Level 5 = "trusted to handle unsupervised and to coach peers." Level 1 = "fails to perform; requires intervention." Levels 2–4 are the intermediate behavioral stages. ### 2.1 Empathy / Emotional Attunement | Lvl | Label | Anchor (observable in transcript) | |---|---|---| | 1 | Fail | No acknowledgement of emotion; jumps straight to policy/transactional response. Customer feels unheard. | | 2 | Advanced Beginner | Cites a scripted empathy line ("I understand your frustration") but moves on mechanically; no follow-up. | | 3 | Competent | Names the emotion in own words, validates it, then transitions to resolution. Appropriate but not tailored. | | 4 | Proficient | Adjusts tone to customer's emotional state mid-call; reflects back specifics ("cracked on arrival — that's frustrating"). | | 5 | Mastery / Entrustable | Reads shifting emotional cues across the call; de-escalates implicitly through pacing and acknowledgment; could model this for new hires. | ### 2.2 Resolution Concreteness | Lvl | Label | Anchor | |---|---|---| | 1 | Fail | Vague ("we'll look into it") or no resolution offered; customer left without a path. | | 2 | Advanced Beginner | Offers a resolution but missing key specifics (no timeline, no method, no amount). | | 3 | Competent | Offers a concrete resolution with method (refund/replacement), amount/channel, and next step. | | 4 | Proficient | Offers a *decision-tree* of concrete options matched to the customer's stated preference; confirms acceptance. | | 5 | Mastery / Entrustable | Tailors resolution to policy + customer constraint, names the exception/risk considered, and closes the loop with a verification step. | ### 2.3 De-escalation | Lvl | Label | Anchor | |---|---|---| | 1 | Fail | Defensive, blames customer/company policy, or matches the customer's escalation. | | 2 | Advanced Beginner | Avoids escalation but through avoidance/deflection rather than active de-escalation. | | 3 | Competent | Uses an explicit de-escalation move (acknowledge → reframe → offer), one cycle. | | 4 | Proficient | Cycles through acknowledge/reframe as needed; lowers intensity without conceding policy inappropriately. | | 5 | Mastery / Entrustable | Prevents re-escalation by reading early signals; preserves relationship and policy simultaneously. | ### 2.4 Professionalism / Conduct | Lvl | Label | Anchor | |---|---|---| | 1 | Fail | Unprofessional language, breaks role, gives prohibited advice (legal/medical/financial), or insults customer. | | 2 | Advanced Beginner | Mostly professional but uses jargon ("RMA", "SLA") or breaks tone once. | | 3 | Competent | Plain-language, in-role throughout, no prohibited advice. | | 4 | Proficient | Adapts register to customer; concise for voice (1–3 sentences); manages silence well. | | 5 | Mastery / Entrustable | Consistently concise, on-brand, voice-appropriate; could serve as a call-center exemplar. | ### Note on anchor design (confidence: **0.78**) - Anchors must describe **observable behavior in the transcript**, not internal states (per good-rubric principles: Jonsson & Svingby 2007; Reddy & Andrade 2010). - Level 3 ("Competent") should be the *passing threshold* and defined as "what a competent entry-level hire would do unsupervised." This makes the 3.5 mean gate (D-032) interpretable as "averaging between Competent and Proficient." - Avoid **evasion anchors** ("somewhat", "mostly") — they destroy inter-rater reliability (Wolfe & Chiu 1997; Barkaoui 2010). The anchors above are behavior-specific. --- ## 3. Mastery Gate N Defensibility (D-032: N=3) ### What the literature says about N-of-M mastery gates - **Bloom (1968) / Block (1971):** Mastery learning classically requires one demonstration at ≥80% but with *corrective instruction between attempts*. The "N" is not the central variable — the *remediation loop* is. Bloom's evidence is on gain, not on N. - **Mastery learning meta-analyses (Kulik, Kulik & Bangert-Drowns 1990; Guskey 2007):** Effect sizes are large (~0.5–0.7 SD) but studies use N=1 with remediation; little direct evidence on N≥2. - **CBME / EPAs (ten Cate 2015; ten Cate & Chen 2018):** Entrustment decisions for an EPA typically require **multiple observations across contexts**. Common recommendations: - **5–10 observations** per EPA is a frequently cited minimum for *high-stakes* entrustment (e.g., surgical EPAs, Rekman et al. 2016). - The ACGME milestone framework treats low-stakes formative entrustment at N=1–2; high-stakes summative at N≥5 with multiple assessors. - **Generalizability theory (Crossley et al. 2002; Bloch & Bogo 2007):** For performance assessments, a single observation has low generalizability (G-coefficients often 0.5–0.7). Generalizability improves with **both** more scenarios *and* more assessors. For voice role-play with one AI assessor, the *scenario count* carries essentially all the reliability burden. - **Standard setting (Norcini & Guille 2002; Cusimano 2014):** High-stakes credentialing exams typically use multi-stage blueprints sampling **multiple content domains** — 3 is on the low end; 6–12 is common for high-stakes OSCEs (Pell et al. 2010). - **Angoff / Ebel methods:** Not directly about N, but the standard-setting tradition implies you sample enough items (scenarios) to cover the blueprint reliably. 3 is thin blueprint coverage. ### Is N=3 defensible? (confidence: **0.62**) **Defensible as a formative / low-stakes gate; not defensible as a high-stakes credential on its own.** Arguments for N=3: - Praxis v0.3 is positioning a "path" credential, not a license to practice. If the credential is employer-facing *internal advancement* (not regulatory), N=3 across *distinct* scenarios satisfies the CBE principle of "demonstrated across contexts" weakly but coherently. - Distinctiveness requirement (D-032 says "distinct scenarios") is the right lever — it's the breadth, not the raw count, that addresses generalizability. Arguments against N=3 (for high-stakes): - A single AI assessor means rater variance is not averaged out; all reliability rides on scenario sampling. G-theory suggests N=3 yields G ≈ 0.5–0.6 — below the 0.8 conventional threshold for high-stakes decisions (Brennan 2001). - 3 scenarios barely covers a blueprint (refund + complaint + escalation = 3 nodes). Real CS skill has more sub-domains. ### Recommended posture (confidence: **0.70**) 1. **Label the v0.3 credential explicitly as "formative" or "path completion"** — not "certification." This makes N=3 defensible. 2. **Add a "high-stakes" tier at N=5–6 distinct scenarios** with blueprint coverage required (≥1 per sub-skill cluster) as the defensible high-stakes threshold. Cite CBME/EPA literature (Rekman 2016; ten Cate 2018) and G-theory (Crossley 2002). 3. **Keep the remediation loop** between attempts — that's where Bloom's mastery-learning effect actually lives. N=3 *without* remediation is weaker than N=1 *with* remediation. 4. **Raise the mean rubric gate from 3.5 to ≥3.5 on each scenario, not just the path mean**, if high-stakes. A path mean of 3.5 can hide a single failing scenario (e.g., 5, 5, 2 → mean 4.0). See §4 for the additive-vs-gating question. 5. Track observed rater-Drift of the LLM extractor over time (D-038); if inter-scenario correlations collapse, N must rise. --- ## 4. Mastery Score Computation ### 4.1 How to combine criteria → scenario score Options: - **(a) Weighted mean of criterion scores** (D-039 has per-skill weights). - **(b) Conjunctive / min-rule** — pass only if *every* criterion ≥ threshold (common in CBME milestone systems; ACGME uses conjunctive for this reason — "no criterion unaddressed"). - **(c) Compensatory mean** — high scores compensate low (what weighted mean implies). - **(d) Hybrid** — minimum floor on critical criteria + weighted mean for the rest (used in many medical licensing rubrics, e.g., MRCP clinical exam). **Recommendation (confidence: 0.74):** Use **(d) hybrid: weighted mean with a floor on critical criteria.** Specifically: - Compute weighted mean of criterion scores (1–5) using D-039 per-skill weights. - Apply a **floor**: scenario passes only if *every* criterion scored ≥ 2 AND the weighted mean ≥ 3.0 (D-032 sets ≥ 3.5 at the path level). - Rationale: A learner who scores 5 on resolution and 1 on professionalism should *not* pass a refund scenario — the floor catches this. The literature strongly favors conjunctive rules for *safety-critical* dimensions (Norcini 2003; Wass et al. 2001 on OSCEs); a hybrid is a pragmatic compromise between conjunctive strictness and compensatory flexibility. ### 4.2 How to combine scenario scores → path Mastery Score **Additive vs gating — the answer is *both*, at different layers.** - **Gating layer (qualitative):** The N-of-M distinct-scenario pass requirement (D-032) is a **gate**, not a sum. You must pass each of N distinct scenarios. This satisfies the "varied-context mastery" requirement from CBME/EPA literature (ten Cate 2018 — entrustment requires demonstrated generalization). - **Additive layer (quantitative Mastery Score):** On top of the gate, compute a numeric Mastery Score as the **weighted mean of scenario scores**, where scenario weights reflect blueprint importance (e.g., harder scenarios weighted higher). This gives a continuous signal for ranking/cohort comparison and for the "rubric mean ≥ 3.5" gate in D-032. **Specific formula recommendation (confidence: 0.72):** ``` MasteryScore(path) = Σ_s ( w_s · ScenarioScore_s ) / Σ_s w_s where ScenarioScore_s = Σ_c ( w_c · CriterionScore_{s,c} ) / Σ_c w_c subject to floor: ∀c, CriterionScore_{s,c} ≥ 2 pass s ⇔ ScenarioScore_s ≥ 3.0 (scenario pass threshold) pass path ⇔ (≥3 distinct scenarios passed) ∧ (MasteryScore ≥ 3.5) ``` This satisfies D-032 exactly: the rubric mean ≥ 3.5 is computed on the *passing* scenarios only (otherwise failed scenarios would drag down a credential earned by passing 3 distinct ones). Decide and document whether MasteryScore is computed over (a) all attempted scenarios or (b) only passing scenarios — **recommend (b)** to align with "mastery" semantics. ### 4.3 Why not just sum? A sum (e.g., "passed 3 of 5 scenarios") loses information about *how well* and creates a perverse incentive to attempt many easy scenarios. The gate + weighted-mean hybrid avoids this. --- ## 5. Deterministic Scoring Patterns (D-038: LLM extracts, rules score) The core problem: free-form speech → reproducible score. The D-038 split (LLM-extracts-evidence, rules-score-evidence) is well-aligned with the literature on **structured rubric scoring from natural language**. ### 5.1 The pattern Two-stage pipelines are the documented way to control LLM variability in assessment (Latif & Zhai 2024 on LLM-as-judge; Chiang & Lee 2023 on explanation-first prompting): 1. **Extraction stage (LLM, allowed to vary):** The LLM is constrained to *extract evidence* — verbatim quotes + structured tags — not to score. Output is a JSON/structured record like: ``` { "criterion": "empathy", "evidence_quotes": ["I'm sorry the item arrived cracked — that's frustrating."], "evidence_signals": ["named_emotion", "acknowledged_specific", "no_policy_first"], "absence_signals": [] } ``` Key: the LLM does **not** emit a number. It emits *what it observed*. This is the documented "evidence-centered design" pattern (Mislevy, Steinberg & Almond 2003) and matches D-038. 2. **Scoring stage (deterministic rules):** A rule function maps `evidence_signals` (+ absence) to a level 1–5 per criterion, per a published lookup table embedded in `rubrics/.yaml`. Identical input → identical output. No LLM in this stage. ### 5.2 Why this beats "LLM scores directly" - **Reproducibility:** Same transcript + same extraction prompt → same evidence tags (modulo LLM nondeterminism, mitigated by temperature=0 + structured output / JSON schema). Rule scoring is fully deterministic given the tags. - **Auditable:** A learner can see *which quote triggered which signal → which level*. This satisfies CBE transparency principles (C-BEN 2023) and is essential for appeals. - **Calibratable:** The signal→level table is editable in YAML without retraining; rubric revision is a config change, not a model change. - **Lower hallucination surface:** LLM is asked only to quote + tag, not to *judge*. Quoting grounds it in the transcript (reduces drift). ### 5.3 Concrete signal taxonomy for one criterion (empathy) ```yaml # rubrics/customer_service.yaml — fragment criteria: empathy: weight: 0.30 signals: - id: no_acknowledgement # absence signal weight: -2 - id: scripted_empathy_line # "I understand your frustration" weight: +1 - id: named_emotion_in_own_words weight: +1 - id: acknowledged_specific # references the actual situation weight: +1 - id: tone_pace_adjusted # extracted from sentence length / hedging weight: +1 - id: policy_first_before_emotion weight: -2 levels: 1: { if: [no_acknowledgement, OR, policy_first_before_emotion], score: 1 } 2: { if: [scripted_empathy_line, AND, NOT named_emotion_in_own_words], score: 2 } 3: { if: [named_emotion_in_own_words, AND, acknowledged_specific], score: 3 } 4: { if: [3-level signals, AND, tone_pace_adjusted], score: 4 } 5: { if: [4-level signals, AND, no_policy_first_before_emotion, AND, >=2 acknowledgement instances], score: 5 } ``` The rule engine evaluates these deterministically. The LLM's only job is to populate the `signals` list with quotes. ### 5.4 Remaining risks and mitigations (confidence: 0.68) | Risk | Mitigation | |---|---| | LLM extraction nondeterminism | temperature=0, fixed seed, JSON schema-validated output, retry-on-schema-fail. | | LLM misses evidence (false negative) | Run extraction twice on borderline cases; flag disagreement for human review. | | LLM tags a signal that isn't in the transcript (hallucinated quote) | Validate that each `evidence_quote` is a fuzzy-match substring of the transcript; reject otherwise. | | Rubric drift across model upgrades | Pin extractor model version (already D-020-style); re-run a golden transcript regression suite on any model change. | | Adversarial phrasing | The signal taxonomy is behavioral; a learner who says the magic words without behavior still lacks the *specificity* and *tone_pace* signals, capping at level 2–3. | **Overall confidence in the two-stage pattern: 0.80** — this is the strongest-evidence recommendation in this document; the extraction/scoring split is well-grounded (Mislevy ECD; Latif & Zhai 2024 survey). --- ## 6. Customer Service Skill Weights (refund/complaint scenario) ### 6.1 Evidence on what matters in CS calls - **Customer satisfaction (CSAT) literature:** Empathy and "soft" dimensions dominate CSAT variance in complaint/refund contexts (Verleye 2004; Makavana 2021 survey of CSAT drivers). Resolution matters but is *table stakes* — customers don't reward it, they punish its absence. - **Service recovery paradox (Magnini, Ford, Markowski & Honeycutt 2007):** After a service failure, *recovery quality* (empathy + ownership) drives loyalty more than the refund itself. This argues empathy ≥ resolution in a *complaint* context specifically. - **De-escalation** is the safety-critical dimension in escalated calls — it prevents churn, legal escalation, and reputational damage. In *non-escalated* calls it's nearly irrelevant. Weight should be context-dependent. - **Professionalism / conduct** is a *floor* dimension, not a weighting dimension — it's the conjunctive floor from §4.1, not something to up-weight. ### 6.2 Recommended weights for a refund/complaint scenario (confidence: 0.70) | Criterion | Weight | Rationale | |---|---|---| | Empathy / emotional attunement | **0.35** | Dominant driver of CSAT in service-recovery contexts (Verleye 2004; service recovery paradox literature). | | Resolution concreteness | **0.30** | Table-stakes; customers punish absence but don't proportionally reward presence. Still substantial because a great empathic call with no resolution is a failure. | | De-escalation | **0.20** | Safety-critical but only activates in escalated branches. Lower default weight because in the *non-escalated* branch it's near-saturated; *raises* in scenarios with an `escalates_unresolved` failure mode (D-009). | | Professionalism / conduct | **0.15** | Treated as floor (conjunctive ≥2 to pass) rather than primary weight. | **Important nuance:** These weights are for the **refund/complaint** scenario specifically (the v0.1 scenario `cs_refund_ca_v01`). A different scenario archetype (e.g., "general inquiry") would tilt empathy down and resolution up. D-039's per-skill weights should be **per-scenario-archetype**, not one global CS weight set. Recommend D-039 be amended to allow `rubrics/customer_service_.yaml` or a weights override block in the scenario file. ### 6.3 Dynamic weighting suggestion (confidence: 0.55 — lower, speculative) If a branch escalates (D-009 `escalates_unresolved` triggered), re-weight on the fly: de-escalation → 0.40, empathy → 0.30, resolution → 0.20, professionalism → 0.10. The rubric's *relevance* changes once the call has gone bad. This is consistent with context-sensitive rubric weighting in OSCE station design (Pell et al. 2010). --- ## Summary confidence table | Section | Confidence | Driver | |---|---|---| | 1. Rubric models (Dreyfus+Miller+EPA+Bloom mastery) | 0.82 | Strong framework fit; well-established literature. | | 2. 5-level anchoring example | 0.78 | Based on established good-rubric principles; example is illustrative, not validated. | | 3. N=3 defensibility | 0.62 | N=3 defensible only for formative / path-completion credentials; thin for high-stakes. | | 4. Mastery score computation (hybrid floor + weighted mean, gate+additive layered) | 0.72 | Aligns with CBE/EPA practice; specific formula is a synthesis, not a direct citation. | | 5. Deterministic scoring (LLM-extract + rule-score) | 0.80 | Strongest evidence base (ECD, LLM-as-judge surveys); pattern is well-grounded. | | 6. CS weights for refund/complaint | 0.70 | Anchored in CSAT/service-recovery literature; specific numbers are judgment calls. | ## Key references - Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). *A Taxonomy for Learning, Teaching, and Assessing.* Bloom's revised taxonomy. - Barkaoui, K. (2010). Do ESL essay raters' evaluation criteria change with experience? *Assessing Writing.* - Benner, P. (1982). From novice to expert. *AJN.* (Dreyfus applied to nursing.) - Block, J. H. (1971). *Mastery Learning: Theory and Practice.* - Bloom, B. S. (1968). Learning for mastery. - Brennan, R. L. (2001). *Generalizability Theory.* (G-coefficient thresholds.) - C-BEN (2023). Quality Assurance Principles for CBE programs. - Chiang, C.-H., & Lee, H.-Y. (2023). Can large language models be good judges? - Crossley, J., Davies, H., Humphris, G., & Jolly, B. (2002). Generalisability in healthcare assessments. - Cusimano, M. D. (2014). Standard setting in medical education. - Dreyfus, H., & Dreyfus, S. (1986). *Mind Over Machine.* (Five-stage skill acquisition.) - Guskey, T. R. (2007). Closing achievement gaps: Revisiting mastery learning. - Jonsson, A., & Svingby, G. (2007). The use of scoring rubrics: Reliability, validity, and educational consequences. - Kulik, C.-L. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). Effectiveness of mastery learning programs. - Latif, S., & Zhai, X. (2024). A systematic review of LLM-as-a-judge. - Magnini, V. P., Ford, J. B., Markowski, E. P., & Honeycutt, E. D. (2007). The service recovery paradox. - Miller, G. E. (1990). The assessment of clinical skills/competence/performance. *Academic Medicine.* - Mislevy, R. J., Steinberg, L. S., & Almond, R. A. (2003). On the structure of educational assessments. (Evidence-centered design.) - Norcini, J. (2003). ABC of learning and teaching in medicine: Work based assessment. - Norcini, J., & Guille, R. (2002). Standard setting in medical education. - Pell, G., Boursicot, K., & Roberts, T. (2010). Could OSCEs be replaced? (Blueprint coverage / station counts.) - Rekman, J., Hamstra, S. J., et al. (2016). Entrustable professional activities. (N recommendations.) - Reddy, Y. M., & Andrade, H. (2010). A review of rubric use in higher education. - ten Cate, O. (2005). Entrustable professional activities. - ten Cate, O., & Chen, H. C. (2018). The EPAs of competency-based medical education. - Verleye, K. (2004). Empathy in customer service. - Wass, V., Van der Vleuten, C., Shatzer, J., & Jones, R. (2001). Assessment of clinical competence.