Blog
The Solvable Paradox — Part 2

Dispatches from D47K

November 18, 2025

Posts and notes on trust, blockchains, and the messy history in between.

ConsensusReputationIncentives

A new consensus approach

In Part 1 I listed how incentives bent “decentralized” systems back toward gatekeepers. This part sketches a path out: True Proof of Work (TPoW) that rewards human contributions, reviewer-driven validation, and reputation that isn’t just a bag you can buy.

I started exploring this while building a crowdsourcing tool: could we reach consensus by rewarding meaningful work instead of hash waste, and do it without creating a new cartel?

True Proof of Work (TPoW)

As stated in Part 1, current PoW burns massive power to run a lottery. PoS swaps watts for wealth. TPoW keeps the randomness for liveness, but pays out based on human contributions that add value, not hash rates or bag size.

Valuable contributions: work that can’t be automated cheaply (e.g., curation, review, human checks), scored by quality, reputation, and difficulty. Rewards are split across contributors according to weight, not paid to one jackpot winner. That dampens “buy better odds” behavior.

Economic guardrails: define who funds the pool (fees, limited issuance, or an external budget), cap total emission, and make rewards decay over time so early contributors don’t lock in perpetual dominance.

Anti-Sybil: joining the reviewer/contributor set requires a bond or a rate-limited identity gate plus reputation aging. Throwaway keys exist, but without reputation and without a bond they earn nothing.

Illustrative payout split (no jackpot)
Contributor Weight Payout share
A (top) 30% 24%
B 25% 20%
C 20% 16%
D 15% 12%
Long tail 10% 28%
Diminishing returns push excess reward to the tail; no single winner takes all.

By the time I started to write this article, together with some friends I had, without realising it, been working on a building block of such a TPoW solution for over one year. The TPoW principle was built into a traditional Software as a Service solution, a crowdsourcing platform for data quality control. We are currently exploring whether we could integrate TPoW in a new, reinvented blockchain protocol.

The prototype we built was a browser plugin that lets internet users report issues they find when browsing the web. Think of it as a generic “report an error” button: it captures the URL and a short note, nothing fancy.

A second prototype on the same engine was a city-issues reporter: potholes, broken traffic lights, vandalism, trash—quick submissions with location and a note.

The principle of the Mechanical Turk is used in both tools, leveraging the power of the crowd to spot and report problems that cannot be identified by software-powered bots. The difference between existing quality control and bug-hunting solutions is that the first step is taken by the user instead of by the owner of the process or platform behind the reported issue.

In traditional systems, the owner starts a bug bounty campaign for instance, where a handful of experts are hired focussing mostly on major issues. Or expensive surveys are conducted to ask people for their opinion. With these prototypes, the user takes the initiative, and the owner of the process is informed after an issue has been reported on the platform.

The contributors who reported issues in a certain time frame (= blocks) then all receive a reward according to the weight (or quality) of their contribution. To keep distribution fair, the number of contributions will only have a minimal impact. A contribution score taking also other elements into account, such as the user’s reputation, the quality of their contributions, the level of a user, the age of a user’s public/private key pair, the number of rejected and accepted issues, feedback scores, etc. will all be used by the algorithm to calculate the user’s Reward Weight for a certain block. Each user then receives a reward score between 0 and 100 and will be rewarded according to his score. Having a maximum per key pair discourages trying to game the system. So, instead of awarding one user as is done in PoW and PoS, all contributing users receive part of the reward and fees of a block.

The Prisoner’s Dilemma

Consensus is then reached by randomly assigning a winner amongst the users who have a high-enough Reward Score (which is any score greater than a pre-defined threshold) and have reached Reviewer status. The minimum Reward Score is fixed and hard-coded in the protocol.

Without a controlling mechanism that safeguards the quality of the contributions made by users, such a platform would be useless. In a traditional blockchain solution, confirmations are done by miners. In the HB and TSCA solutions, we replaced the miners with a network of Reviewers.

A network of nodes then:

1. keeps a copy of the ledger with reported issues;

2. keeps a record of recent issues that have been reported that still need reviewing;

3. keeps track of balances and monetary transactions like in a traditional blockchain.

Every newly reported issue is randomly assigned to a minimum of five Reviewers. These five Reviewers need to reach a consensus on whether an issue is a valid contribution or not. A Reviewer only receives one new issue at a time. The Reviewer can then choose to accept or reject the issue (after inspection), or an issue can be skipped.

Skipping too many issues will cause the Reviewer to lose Reputation points. A negatively impacted Reputation Score will inevitably result in the Reviewer losing his Reviewer status. The same happens when a Reviewer chooses the wrong option when accepting or rejecting an issue. This will filter out Reviewers trying to game the system. Random assignment, hidden votes, and rotating committees raise the cost of collusion.

Reviewer committee (simplified)
  1. Randomly sample reviewers from eligible pool (reputation + bond).
  2. Assign one issue per reviewer; votes stay hidden until finalized.
  3. Majority threshold >60%; else add 2 reviewers (max 9) then timeout.
  4. Slash outliers who repeatedly land on the wrong side of honest majorities.
  5. Rotate committees to prevent static coalitions.
Hidden votes + rotation blunt bribery; slashing outliers makes “spray and pray” expensive.

An example:

Reviewer votes (example)
Issue Bob Claire John Sidney Mary
1 Accept Accept Accept Accept Accept
2 Accept Accept Reject Accept Accept
3 Reject Reject Accept Reject Reject
4 Accept Skip Accept Skip Accept
5 Reject Skip Accept Reject Reject
Votes stay hidden until finalize; the mix shows how skips and dissent trigger penalties.

As we can see in the example, various issues have been treated by five different Reviewers. For the sake of the example, we pretend that all issues were reviewed by the same five Reviewers: Bob, Claire, John, Sidney, and Mary (in real-life situations, this will not be the case though). For every issue, different Reviewers will be randomly assigned. It is important to understand that due to randomisation it is very unlikely that the Reviewers actually know each other. Also, the identity of other Reviewers, nor the option they chose are never shown or communicated to their peers. This way each one’s vote is anonymous and unbiased.

In our example, there are two Reviewers who always have the same vote as the majority of Reviewers. Those Reviewers are Bob and Mary. The other three sometimes chose a different option. Claire skipped two issues, John chose the opposite option in three out of five issues and Sidney skipped one issue:

Reputation impact (example)
Reviewer Start rep Majority aligns Skips Wrong calls Net change End rep
Bob 72.0 +0.5 0 0 +0.5 72.5
Claire 71.0 +0.3 -0.4 0 -0.1 70.9
John 71.0 +0.2 0 -1.5 -1.3 69.7
Sidney 70.5 +0.3 -0.2 0 +0.1 70.6
Mary 73.0 +0.5 0 0 +0.5 73.5
Align with honest majority to climb; skips and wrong calls drag you down fast.

At all times, we want to safeguard the quality of the work of the Reviewers, to avoid any of them gaming the system. To become a Reviewer, one should first participate in the network as a normal user, building up a high Reputation Score. To obtain this Reputation Score, one would have to do valuable contributions which require intellectual labour.

Let’s say that a new user starts with a Reputation Score of 50 / 100. Each time a contribution made by a user was accepted by the randomly assigned Reviewers, the Reputation Score of the user goes up by +1 Reputation Point. To keep our users honest, there will be a fault tolerance which negatively impacts the user’s Reputation Score with -3 Reputation Points. Users with a score lower than 0 will be ignored by the system and cannot participate anymore. A user can at all times create a new key pair and start over.

Reviewing an issue has a similar effect on a Reviewer’s Reputation Score, though with stricter fault tolerance. Being part of the majority awards the Reviewer with +0.1, whereas a negative score costs the Reviewer -0.5 Reputation Points. Skipping an issue costs -0.2 Reputation Points. Any user can become a Reviewer by doing a certain number of contributions and by obtaining a Reputation Score of 70 / 100 (for instance).

Let’s look at our example to see how this impacts the score of each Reviewer:

Reputation math (per issue, illustrative)
Issue Majority? Bob Claire John Sidney Mary
1 Yes +0.1 +0.1 +0.1 +0.1 +0.1
2 Yes +0.1 +0.1 -0.5 +0.1 +0.1
3 Yes -0.5 -0.5 +0.1 -0.5 -0.5
4 Yes +0.1 -0.2 (skip) +0.1 -0.2 (skip) +0.1
5 Yes -0.5 -0.2 (skip) +0.1 -0.5 -0.5
Majority alignment is steady lift; skips and wrong calls cut faster. One bad round can cost your reviewer seat.

We clearly see the effect of fault tolerance in this example:

§ For Claire, who had three correct answers, but skipped two, the result is negative. She also has to be careful not to lose her Reviewer status.

§ Sidney had one issue that was skipped, which had a minimal impact on the Reputation Score. With the four other answers being correct, the result is still positive after five questions.

§ John was clearly trying to game the system, with three wrong answers out of five. The result over five questions is a negative score of -1.3 which is subtracted from his total. His new Reputation Score is below 70 and he lost his Reviewer status and access to new cases is denied to him.

By limiting the number of times a user can reach Reviewer status to 3 for instance, the stakes are high to make sure that the answers are correct. After all, a Reviewer is an expert in what they do, so a Reviewer’s Reputation Score falling below the threshold is not normal behaviour. Reputation should also decay over time, so inactive keys don’t sit atop the pile forever.

This system is a variation of the Prisoner’s Dilemma, which is used in game theory to show what can happen if two players don’t cooperate with each other. The theory came into existence in 1950 and was baptised Prisoner’s Dilemma after the example used to prove the theory.

Thanks to the Prisoner’s Dilemma we know that users will operate in their own interest when they don’t know what option the others chose, nor do they know the outcome of the group vote. All they know is that if they provide a wrong answer, this will most likely have a negative impact on them. The theory assumes that players understand the consequences of their actions (= they know the rules), the players have no loyalty to each other, and there is no opportunity for retribution or reward outside of the game. To handle real-world bribery attempts, keep votes hidden until finalized, monitor correlated voting patterns, and slash outliers who keep landing on the wrong side of honest majorities.

Attack sketch (bribe vs. defenses)
Step Attacker move Mitigation
1 Offers bribes to reviewers Votes hidden until finalize; no pre-commit info
2 Targets known addresses Random committee; frequent rotation; address churn
3 Pays for wrong votes Slash outliers; decay reputation; bond loss > bribe
4 Attempts repeated rounds Timeouts; extra reviewers; cost grows each round

With the proposed interactions between users and Reviewers, and amongst Reviewers themselves, the rules are very clear:

§ A majority vote strengthens the Reviewer’s own position

§ The users and Reviewers don’t know each other, and their identities will never be revealed to each other, eliminating pre-existing bias

§ Anonymous participation makes peer-to-peer rewards or retribution impossible

The Prisoner’s Dilemma has been heavily researched, and results show that in games that are played, users will choose to act out of self-interest, rather than to game the system since that will always have a negative impact on the user’s own position. The theory has proven its worth in many other domains, such as business, government, sports, economics, and so on.

Going back to our example, you might have noticed that the outcome of Issues 4 and 5 is unknown. To build in some extra safety, consensus can only be reached if the majority vote is higher than 60% of the votes. In case of a majority vote with a value lower than 60% but higher than 50%, extra Reviewers will be asked to confirm or reject the issue. Latency guardrail: cap the number of rounds and define a timeout so blocks don’t stall forever.

Fault tolerance (example)
Round Reviewers Threshold Action
1 5 >60% (4/5) If <60%, add 2 reviewers
2 7 >60% (5/7) If <60%, add 2 reviewers
3 9 >60% (6/9) If <60%, mark unsolvable; no penalties
Timeouts cap block latency; disputes that stay murky get dropped, not dragged.

Up to two times, two extra Reviewers (no more, no less) will automatically be added in case of a majority score lower than or equal to 60%. This implies that for a normal scoring situation, five Reviewers are sufficient. If the majority vote equals 3/5 though, two more Reviewers are asked for their input. If the majority vote, then still didn’t climb above 60% (= a score of 4/7), two more Reviewers are added. If they don’t reach a consensus with a majority of more than 60% (which would mean a score of 5/9), the issue is marked as unsolvable. Discarded issues cannot have a negative impact on a user’s Reputation Score. All input is ignored. No extra Reviewers are added once the maximum of nine Reviewers has been reached.

The TPoW algorithm describes a ledger that stores non-financial information and only takes the distribution of fees and new block rewards into account from a financial point of view. Once coins have come into circulation, transactions with those circulating coins will also need to be written to the mempool to be processed.

Reviewers then replace miners; the mining software is integrated into the widgets used for reviewing non-blockchain-related issues. The same consensus principles as in PoS with a randomly chosen winner can be used then. To reward valuable contributors, the winner could (for example) be randomly chosen from the pool of top performing Reviewers during the previous block. Traditional blockchain mining then becomes a process running in the back of a normal software solution.

Every winner could be rewarded with a token that holds no direct financial value, such as a governance token or an NFT. In case a governance token is used, a maximum of for example five such tokens can be won by a single Reviewer. If governance is attached to such tokens, then they should decay over time, for instance with 0.1 per x number of blocks. New tokens can then only be won once a Reviewer’s reward token balance drops below 4. It is important to not distribute any tokens or coins with a monetary value though, as this would encourage gaming the system with potential centralisation as an unwanted indirect result. If you do need monetary rewards, back them with a finite pool (fees/issuance), add slashing, and cap per-key earnings with diminishing returns.

Anonymous reputation building

Using a user- and task-driven consensus mechanism opens up opportunities in other areas as well. One of those is reputation building. Bitcoin came into existence as a pseudonymous coin. Later on, privacy coins like Monero were launched, aiming to provide its users with full anonymity.

Linking anonymity to financial transactions with cryptocurrencies or with smart contracts left us with some new challenges, affecting all parties involved. Peer-to-peer transactions come with a risk since we don’t always know who we’re dealing with and whether the other party can be trusted. In case of problems, there’s no bank or other trusted party anymore where we can complain to get our money back. E-commerce platforms and financial tools and websites on the other hand are forced to do a full KYC (Know Your Customer) and AML (Anti-Money Laundering) of every client, and in some countries deposit addresses for crypto need to be linked to the user’s profile.

Two worlds collide here, as such an approach tries to solve a new-tech problem with old-school, low-tech approaches. This works counterproductive and significantly slows down onboarding processes. And then we haven’t even mentioned the fact that different websites and apps often use different KYC providers.

The result is that the user needs to go through KYC over and over again. Personal documents such as identity cards, passports, utility bills, tax return forms, bank statements, etc. have to be shared with KYC companies that have no further affiliation with the user. Plus, the user instantaneously loses oversight of who holds what data about them, nor does the user know how this is stored and secured, how many backups there are and for how long data is kept.

These issues could be avoided by using the power of a user-driven blockchain system, where participants build up a reputation when they interact with others.

Reputation scoring

In real life, one’s name ends up in various databases and many of those systems give their users some sort of reputation score: how much do you spend on average, how many orders do you place per month or per year, how fast do you pay your bills, are you responsive, are you in debt, etc. The disadvantage is that a negative score will likely stick to the user. If someone goes bankrupt for instance, then this will always be visible in some databases, which undermines ones right to get up and try again.

If we would use a reputation-building blockchain solution, then everyone could anonymously create as many public/private key pairs as they want. These key pairs would then exist on the blockchain with a reputation score linked to it. Providing another person or a company with a public key will then be sufficient to prove one’s reputation or credibility.

Of course, when running such a system, all possibilities to game the system should be taken out of the protocol. As I already described above in the chapter about the Prisoner’s Dilemma, every key pair would then have to start with a neutral score of 50 / 100, with 0 being the lowest possible score and 100 being the highest score one could get.

By applying TPOW tasks, users can build up a reputation score in a certain area of expertise for instance. But also suppliers, such as app or website developers or other third parties would receive a reputation score. This is already done on many platforms by letting users give a number of stars or points to a certain supplier. Various systems are being used though and they don’t allow benchmarking, although this could be beneficial for the user of such platforms. Also, we’re never sure whether the supplier cheated by letting employees, friends and family score their own platform multiple times, thus creating a false impression.

Let’s assume a new user signs up for a key pair. The user then receives two scores, right from the start: a neutral reputation score of 50/100 and a level of 0/10, where 0 is an absolute beginner and 10 is an expert. One could have a level of 8/10 for instance, which means the user has been around for a while already, but still have a Reputation Score of 50/100, which would imply that this experienced user is struggling to build up a good reputation. This puts his score in perspective versus a new user who starts with a score of 50/100.

The scoring system is then a rolling number of points that can go up or down depending on the actions and feedback a user receives, optionally combined with machine learning. Positive feedback could result in 1 point being added to the user’s total. Fault tolerance would then be used for negative feedback, subtracting 3 points from the user’s total. Users with a Reputation Score of 40 or lower would automatically be ignored by the system since they’ve proven to not be trustworthy.

Reputation over time (illustrative)
Block span Active reviewer Inactive key
Start 50/100 50/100
+100 blocks 62/100 (steady positives) 48/100 (decay)
+500 blocks 78/100 40/100 (ignored if below floor)
Reputation ages and decays; staying active and correct matters more than hoarding an old key.

As a user’s reputation builds up slowly by providing valuable contributions as either Hunter, Reviewer, Affiliate, or Client… the willingness of other users to interact with such user will increase as well. The user builds up an anonymous Reputation Score with no need for a trusted party to do extensive background checks. This can be seen as trustless trust.

The effort of slowly having to build up a reputation will outweigh the risk of one’s Reputation Score rapidly tumbling down as a result of taking fraudulent actions. A user’s score will be linked to his key pair, which is visible on the Reputation Blockchain. This blockchain can be queried by anyone using block explorers or chain analysis software. The higher one’s Reputation Score, the higher the chance of others willing to trust and interact with that user. Thanks to bi- or even omnidirectional reputation scoring, all parties involved in a transaction will have to prove their honesty and knowledge in a certain area of expertise.