SESSION_001
Labeled
0
of 8 prompts
Response A Wins
0
—%
Ties
0
—%
Response B Wins
0
—%
A
Annotator 1
TASK_ID: PML-001 · Role: Senior Reviewer
1 / 8
Ann-1
User Prompt
Loading...
Response A
Response B
Quality Dimensions
A Select A
B Select B
T Tie
Submit
S Skip
Recent Annotations
No annotations yet.
Total Annotators
3
active this session
Avg Agreement
Cohen's κ
Conflicts
0
need resolution
Resolved
0
consensus reached
Annotators

+ ADD ANNOTATOR

Inter-Annotator Agreement
Cohen's Kappa Score
Annotate more to compute
Pairwise Agreement Matrix
Annotate prompts with multiple reviewers to see agreement.
⚠ Conflict Queue
No conflicts detected.
✦ AI Prompt Generator
Reasoning
Code
Safety
Math
Creative
Factual Q&A
Ethics
Roleplay
Ready to generate prompts
↑ Upload Prompt Batch
📁
Drop JSON batch file
Format: [{prompt, responses:{a,b}}]
Prompt Queue
8 prompts
Add Prompt Manually
Export Format
🧠
RLHF Training Pairs
chosen/rejected pairs for reward model training
Comparison Dataset
Full preference scores with all dimensions
🎯
DPO Format
Direct preference optimization (prompt, chosen, rejected)
📋
Raw JSON
All annotation data, unprocessed
Preview
{ "message": "No annotations yet. Go to Annotate tab to label some prompts." }
Pipeline Status
1
Prompts Loaded
0 prompts in queue
2
Annotations Collected
0 labeled, 0 skipped
3
Quality Filtered
Run quality check first
4
Format Ready
Select format above
5
Download / Push
Export to training pipeline
Dataset Summary
Total pairs0
A preferred0
B preferred0
Ties excluded0
Annotators1
Overall Quality
weighted score
Consistency
same Q, same answer
Flagged
0
need review
Bias Score
position bias check
Annotator Quality Scores
Submit annotations to see quality scores.
⚠ Flagged Items
No flags raised.
AI Quality Analysis
✦ Run AI Analysis
Claude will analyze your annotation patterns, detect potential biases, and suggest quality improvements.
Position Bias Check
Are annotators preferring A just because it appears first?
chose A
vs
chose B
Annotate at least 5 prompts for bias detection.