How DataBounty works

DataBounty is a platform for open coding datasets. Sponsors request the work, contributors create the items, and validators audit quality. A contract-driven verification pipeline sits in the middle so that the platform reports what has actually completed.

Coding is the live domain today; legal, healthcare, finance, and math/science are opening next. see_all_domains →

Live today: open datasets, rewarded with karma

Submit to open pools directly. In a policy-controlled pool, final acceptance releases karma immediately and queues eligible work for Hugging Face synchronization. New community specs and dataset types open regularly. High-karma contributors earn first access. how_karma_works →

01

The loop

The active community path, from a live pool contract to verified publication.

01 choose_open_pool02 submit_items03 contract_review04 final_acceptance05 shared_review_window06 verified_publication
  1. 01
    Choose an open pool

    Browse the active community pools and read the live contract: field schema, verification profile, remaining capacity, and karma per final accepted item.

  2. 02
    Submit useful items

    There is no claim step for an open pool. Submit within the contract and declared generation method; capacity is checked by the server.

  3. 03
    Validation and review

    Items move through the dataset contract's declared checks. A required check that is unavailable or unresolved routes the item to review rather than being called verified.

  4. 04
    Final acceptance

    For a policy-controlled pool, a validator approval in full-human mode—or a clean required-pipeline result in automation-only mode—is final. Rejected or unresolved items do not qualify.

  5. 05
    Immediate karma

    Policy-controlled pools release karma as soon as the item is finally accepted; there is no sponsor dispute or shared hold window.

  6. 06
    Asynchronous publication

    Each accepted item is queued for Hugging Face synchronization without delaying karma. Published datasets can include contributor credit when eligible and not opted out.

02

The verification model

The contract defines the required checks. A missing or unresolved required check routes to review rather than a verification claim.

layer 1

Duplicate wall

Each submission is similarity-scored against everything already accepted in the dataset, and across the platform. Near-duplicates are rejected before they cost anyone review time.

layer 2

Sandboxed test execution

For execution-verified types (debugging, implementation, SQL, regex, translation, performance) the contract is executable: broken code must fail the submitted tests and fixed code must pass all of them. If the checks do not hold in the sandbox, the item bounces automatically.

layer 3

Contract-aware quality review

When the dataset contract requires an automated or human quality stage, its completed state is recorded with the item. An unavailable or unresolved required stage routes to review; it is never presented as a pass.

layer 4

Human validation (when required)

A policy-controlled community pool is either full-human, where every automation-cleared item is reviewed, or automation-only, where the clean required pipeline is final. The sponsor does not choose coverage or a sampling rate.

AI-assisted creation is allowed

We verify the output rather than the process. Contributors may use AI tools, or generate items entirely with AI, as long as the generation method is disclosed on each submission. Every item faces the same pipeline either way, and each dataset publishes its human / AI-assisted / AI-generated mix so users know exactly what they are getting.

03

Ranks

Ranks are earned per role. Higher ranks unlock capacity and access, but they are never a substitute for passing the pipeline.

</> builder_ranks (contributors)
1 Scout1 active batch
2 Apprentice2 active batches
3 Builder30-item batches
4 Specialist75-item capacity
5 Craftsmanpriority batch access
6 Senior Buildersenior capacity
7 Expertexclusive datasets
8 Architectarchitect capacity
9 Principalspec consultation invites
10 Master Buildertop claim priority
◇ auditor_ranks (validators)
1 Observersmall audit batches
2 Reviewerstandard batches
3 Inspectorstandard batches
4 Auditorlarger + bonus mult.
5 Senior Auditorfull-audit datasets
6 Quality Leadquality-lead access
7 Verifierverifier duties
8 Principal Verifierdispute-panel access
9 Arbiterarbitration duties
10 Master Arbitertop bonus multiplier
04

Common questions

The short version of the rules everyone plays by.

What is available today?

The public launch is the community program. Each open pool shows its own contract and karma per final accepted item.

What does karma mean?

Karma is a reputation record. The exact amount is shown on the pool contract. In a policy-controlled pool it releases on final acceptance; the contract states any other lifecycle explicitly.

How long does publication take?

Hugging Face synchronization runs asynchronously after final acceptance. It does not delay karma in a policy-controlled pool; unresolved required checks still prevent final acceptance.

Is AI-generated work allowed?

Yes. AI assistance and full AI generation are allowed, with disclosure. We verify the output rather than the process: every item faces the same duplicate, execution, and validation checks regardless of how it was made. Each dataset publishes its generation mix.

How is quality described honestly?

The dataset contract names the checks that apply. DataBounty only presents an item as verified after the required stages are completed; otherwise it shows the precise pending, failed, or review state.

Can I work with AI assistance?

Yes, where the pool contract allows it. Declare the generation method honestly; every submission is evaluated against the same contract and evidence requirements.

Ready to see it in motion?

Browse the open datasets or head to the dashboard to start.