DataBounty is a marketplace for coding datasets. Sponsors request and fund them, builders create the items, and validators audit them. An automated verification pipeline sits in the middle so that only work that provably matches the spec gets accepted and rewarded.
Coding is the live domain today; legal, healthcare, finance, and math/science are opening next. see_all_domains →
Claim batches with no bond, earn karma per accepted item, and get named credit when eligible work publishes to Hugging Face. Paid bounties, funded in USDC and paid per accepted item, are coming soon. Karma holders get first claim. how_karma_works →
From spec to delivered dataset in ten steps. Every step is funded, verified, and auditable.
The sponsor specs the dataset: category, language, item count, difficulty mix, license type, and deadline. Nothing is public yet.
The platform drafts a structured plan (slots, per-item rewards, pilot size) and reviews the spec for feasibility before it can be funded.
Roughly 1-2% of the total budget is deposited in USDC. This opens a small pilot batch to trusted contributors, so exposure stays tiny.
The sponsor inspects real accepted pilot items. They can approve, request one spec revision, or walk away. Silence auto-approves after 7 days.
The spec locks and the remaining budget moves to escrow. From here, spec-matching work cannot be rejected on taste.
Contributors claim 10-item batches inside slots, each with a clear reward per item and a 72-hour deadline. No bond while on time.
Every submission runs the automated gauntlet: duplicate wall, benchmark-contamination screen, sandboxed test execution, LLM validation.
Depending on audit mode, ranked validators review some or all provisionally accepted items and flag real issues for bonuses.
Accepted items trigger USDC payouts from escrow to contributor wallets; validators earn base rewards plus confirmed-issue bonuses.
The export is handed to the sponsor under the chosen license, with quality stats and a public sample published on the marketplace.
Five layers between a submission and a payout. An item must clear all of them, with no exceptions and no manual overrides.
Each submission is similarity-scored against everything already accepted in the bounty, and across the marketplace. Near-duplicates are rejected before they cost anyone review time, and they are never paid.
Items are checked against public coding benchmarks and well-known problem sets. A dataset that leaks eval questions is worthless for training, so likely benchmark copies are flagged and blocked.
For execution-verified types (debugging, implementation, SQL, regex, translation, performance) the contract is executable: broken code must fail the submitted tests and fixed code must pass all of them. If the checks do not hold in the sandbox, the item bounces automatically.
A model reviews each item against the bounty spec (prompt clarity, solution correctness, difficulty calibration, explanation quality) and produces a pass score with notes.
Bounties choose LLM-only, partial (25% sample), or full audit. Ranked validators re-review provisionally accepted items and earn a bonus for every confirmed issue, a direct incentive to find real problems.
We verify the output rather than the process. Contributors may use AI tools, or generate items entirely with AI, as long as the generation method is disclosed on each submission. Every item faces the same pipeline either way, and each bounty publishes its human / AI-assisted / AI-generated mix so buyers know exactly what they are getting.
Nobody commits the full budget to an untested spec. The pilot is the cheapest possible way to find out if the spec is right.
Only about 1-2% of the budget is deposited to open the pilot. A small batch is produced by trusted contributors and runs the full pipeline, so the sponsor judges real, verified samples.
The sponsor has 7 days to approve the pilot or request one spec revision. If they do nothing, the pilot auto-approves, so contributors are never left waiting on an absent buyer.
Once the full tranche is funded, the spec is frozen. Work that matches the locked spec and passes the pipeline cannot be rejected. Taste disputes belong in the pilot phase, before contributors have done the work.
Everything settles in USDC on Solana. The fee schedule and escrow rules are the same for every bounty.
The fee covers the validation pipeline, sandbox execution, escrow operations, and dispute resolution. Every delivered dataset can later be listed in the resale catalog; the license only sets how long it is held off the market first: 30 days on Standard, 12 months on Extended. Contributors keep a 50% pro-rata resale royalty. The fee is carved out of the budget up front and shown on every bounty page.
Deposits sit in the platform treasury, reserved per bounty into contributor pool, validator pool, and fee. Rewards accrue as items are accepted and are paid out in periodic USDC batches to contributor and validator wallets.
If a bounty ends below target, every accepted item is still paid in full. The sponsor takes delivery of the partial dataset at a prorated price and the unspent escrow is refunded. Underfilling is a priced outcome, not a dispute.
There is no upfront stake. A contributor or validator who abandons a claimed batch must post a bond (10% of the expected reward) on future claims until their record recovers. Deliver on time and you never lock a token.
Ranks are earned per role. Higher ranks unlock capacity and access, but they are never a substitute for passing the pipeline.
The short version of the rules everyone plays by.
The sponsor: a lab, company, or individual who wants the dataset. They deposit USDC on Solana into the platform escrow in two tranches, a small pilot first, then the full budget after they approve real pilot samples.
Contributors earn a fixed USDC reward per accepted item (set per slot). Validators earn a base reward per audit batch plus a bonus for each confirmed issue they flag. All rates are published on the bounty before anyone claims work.
Partial fill is a first-class outcome. All accepted items are paid at the full per-item rate, the sponsor receives the partial dataset at a prorated price, and the unspent escrow is refunded. Contributors never eat the shortfall.
Yes. AI assistance and full AI generation are allowed, with disclosure. We verify the output rather than the process: every item faces the same duplicate, contamination, execution, and validation checks regardless of how it was made. Each bounty publishes its generation mix.
No. After the pilot is approved, the spec locks. Items that match the locked spec and pass the pipeline are accepted and paid from escrow. The pilot phase exists precisely so taste disagreements get resolved before the full budget is committed.
Browsing, claiming, and small earnings are wallet-only. Once cumulative payouts to a wallet cross the regulatory threshold (currently $600), payouts pause until identity verification is completed. This is shown in your dashboard well before it applies.
Browse the open datasets or head to the dashboard to start.