Annotator Study_
Welcome

Welcome to the Annotator Study

The aim of this study is to test whether an automatic, reference-free quality score can rank annotators by how carefully and accurately they work. Your masks, drawn at different paces, are what we compare that automatic score against.

Twenty-four real pathology nuclei and dermoscopy lesions patches are to be annotated. For each one, you'll draw the boundary with the brush or lasso, no pre-drawn outline to start from. How long this takes is individual and depends on your own experience. Speed and accuracy both score. At the end you will send your results to me via an email button with all the necessary information.

How this study is structured
  1. A few quick personal questions.
  2. More questions of your experience before you start.
  3. What you'll be looking at: two example photos.
  4. Short tutorial and a practice canvas, right before the real task, no scoring.
  5. Block 1: 12 patches, at one pace.
  6. Block 2: the other 12 patches, as long as you need.
  7. A short survey about the task.
  8. More question, for a possible future study.
  9. Your results: press the email button, then attach the downloaded results file and send it.
Annotator Study_
Personal information

Who's at the bench?

What is your professional or academic background?
Years of experience annotating or analyzing biomedical images (histopathology, dermoscopy, or similar)?
Do you have formal training in reading histopathology or dermoscopy images (medical, biology, or similar background)?
How scoring works
Match
Overlap (Dice) between your mask and the real ground truth.
Speed
A bonus that fades out after about 90 seconds on a patch. Careful beats frantic, but don't stall.
Streak
Consecutive clean matches (≥85%) build a streak. A rough one resets it.
Before you begin

What you'll be looking at

Two quick examples of the raw photos in this bench, no markings, just what the tissue/lesion actually looks like. In the real task, the canvas starts blank. Your job is to draw the boundary yourself with the brush or lasso tools, entirely from scratch.

Take your time to look at the texture and edges here. The real patches are the same size and quality as what you see now. When a patch loads for real, look for the true edge of the nucleus or lesion, and trace it as closely as you can.

Right before you start

Example: a finished annotation

Here is what a correctly traced boundary looks like on each of the two photos from before, one for pathology and one for dermoscopy, for reference only, not something to copy exactly since every patch is different. This is the exercise for both domains: an outline that follows the true edge of the object as closely as you can manage.

Try it yourself

Practice on these images before the real task starts, pathology on the left, dermoscopy on the right. Nothing you draw here is scored or submitted, it's only so you're comfortable with the tools first.

Tool
Brush: click and drag along an edge. Lasso: click around a region's outline, release to close and fill it.
Before you start

A few quick questions

No wrong answers, this just helps us interpret the scores afterward.

How familiar are you with cell/nucleus segmentation specifically? (1 = not at all, 5 = very)
Have you used a digital brush/lasso image-editing tool before (Photoshop, GIMP, any annotation software)?
Right now, how alert do you feel? (1 = very tired, 5 = fully alert)
What are you using to draw masks right now?
Do you have any diagnosed color vision deficiency?
Bench closed

Before your results, how was it?

These help us design the scoring/gamification side better for real medical annotation work.

How motivating did the points system feel while you worked? (1 = not at all, 5 = very motivating)
Did the streak indicator change how you worked?
How enjoyable did you find this task overall? (1 = not enjoyable, 5 = very enjoyable)
Would seeing a LIVE leaderboard of other annotators' scores have changed how you worked?
Which tool did you rely on more?
How confident are you in the accuracy of your final masks overall? (1 = not confident, 5 = very confident)
Did you actually work differently between the "careful" block and the "rushed" block?
Which of your two blocks do you think came out more accurate overall?
Did you run into any technical issues while drawing masks (lag, misclicks, accidental submits)? If so, please write it in "Optional:[...]" at the end of this page.
Which type of boundary mistake do you think you were most likely to make yourself while drawing?
Did you typically review your mask before submitting, or submit as soon as it looked roughly right?
Which factors do you think should count most when judging the quality of an annotator's work? (choose all that apply)
Which kind of feedback would help you improve your annotation skills? (choose all that apply)
Which factors motivate you to participate in annotation tasks like this one? (choose all that apply)
Optional: any other thoughts on the task?
One more thing

For a future version of this task

This task didn't use points, badges, or a leaderboard. We're considering them for a separate future study, and your honest preference helps design that, it has no effect on the results you just submitted.

If a future version of this task included game elements, which would you find most engaging? (choose all that apply)
Do you think game elements like points, badges, or a leaderboard might make annotators less careful, for example prioritizing speed or a good score over real accuracy?
How often do you personally play games or use gamified apps (fitness trackers, language apps, and similar)?
Have you used a "game with a purpose" before, where playing also produces useful data or labels (examples: Foldit, Zooniverse, Duolingo)?
If a future version showed live feedback on your score, when would you want to see it?
If a future version had a leaderboard, how should it identify people?
Would game elements make you more likely to volunteer for future annotation tasks like this one?
One more thing

An automated starting mask

We're also considering giving annotators an automated starting mask (an AI-generated pre-segmentation) in a future version, instead of always starting from a blank canvas. This task didn't have that, so these are about a hypothetical future version too.

If a future version gave you an automated starting mask that you could accept, edit, or redraw from scratch, would that help you?
If an automated starting mask were offered, how would you want to use it?
How much would you trust an automated starting mask without carefully checking it yourself?
Do you think relying on an automated starting mask might make annotators less careful, for example accepting it without close inspection?
Block 1 of 2

Work carefully

Take your time. Aim for the most accurate boundary you can manage. There's no time limit.

Annotator-
Mode-
Points0
Streak0
Time left30.0s
Patch 1 / 24
SPEC-000
Tool
Action
Brush size
fine28pxbroad
Not able to annotate on your own? Load a rough automated mask and adjust it instead. This is recorded. Loading it keeps anything you already drew; removing it only takes the automated part back out.
Annotator Study_
Bench closed
0 total points,
Avg match
-
Best streak
-
Total time
-
Badges earned
Specimen by specimen
#SetModeMatchTimeHintPoints
Annotator Study_
Group reveal

Load everyone's results

Each annotator downloads a results file when their run ends. Drop every file you've collected here. Nothing leaves this page, it's all read locally in your browser.

Drop .json result files here, or click to choose