⚖
Preference
ML
Annotate
Annotators
Admin
Export Pipeline
Quality Scores
SESSION_001
↓ Export
Labeled
0
of
8
prompts
Response A Wins
0
—%
Ties
0
—%
Response B Wins
0
—%
A
Annotator 1
TASK_ID: PML-001 · Role: Senior Reviewer
1 / 8
Ann-1
Switch Annotator
User Prompt
Loading...
Response A
○
Response B
○
Quality Dimensions
Annotator Notes
✓ Submit & Next
⇔ It's a Tie
→ Skip
✦ AI Assist
A
Select A
B
Select B
T
Tie
↵
Submit
S
Skip
Recent Annotations
No annotations yet.
Total Annotators
3
active this session
Avg Agreement
—
Cohen's κ
Conflicts
0
need resolution
Resolved
0
consensus reached
Annotators
+ ADD ANNOTATOR
Cyan
Pink
Gold
Violet
Green
Add Annotator
Inter-Annotator Agreement
Cohen's Kappa Score
—
Annotate more to compute
Pairwise Agreement Matrix
Annotate prompts with multiple reviewers to see agreement.
⚠ Conflict Queue
No conflicts detected.
✦ AI Prompt Generator
Topic Areas
Reasoning
Code
Safety
Math
Creative
Factual Q&A
Ethics
Roleplay
Count
Difficulty
Easy
Medium
Hard
Expert
Ready to generate prompts
✦ Generate with Claude AI
↑ Upload Prompt Batch
📁
Drop JSON batch file
Format: [{prompt, responses:{a,b}}]
Or paste JSON directly
Load from Paste
Prompt Queue
8 prompts
Clear All
Add Prompt Manually
Prompt
Response A
Response B
+ Add Prompt
Export Format
🧠
RLHF Training Pairs
chosen/rejected pairs for reward model training
⚖
Comparison Dataset
Full preference scores with all dimensions
🎯
DPO Format
Direct preference optimization (prompt, chosen, rejected)
📋
Raw JSON
All annotation data, unprocessed
Preview
{ "message": "No annotations yet. Go to Annotate tab to label some prompts." }
↓ Download Dataset
⎘ Copy to Clipboard
Pipeline Status
1
Prompts Loaded
0 prompts in queue
2
Annotations Collected
0 labeled, 0 skipped
3
Quality Filtered
Run quality check first
4
Format Ready
Select format above
5
Download / Push
Export to training pipeline
Dataset Summary
Total pairs
0
A preferred
0
B preferred
0
Ties excluded
0
Annotators
1
Overall Quality
—
weighted score
Consistency
—
same Q, same answer
Flagged
0
need review
Bias Score
—
position bias check
Annotator Quality Scores
Submit annotations to see quality scores.
⚠ Flagged Items
No flags raised.
AI Quality Analysis
✦ Run AI Analysis
Claude will analyze your annotation patterns, detect potential biases, and suggest quality improvements.
✦ Analyze with Claude
Analyzing...
Position Bias Check
Are annotators preferring A just because it appears first?
—
chose A
vs
—
chose B
Annotate at least 5 prompts for bias detection.
Modal
✓