← Blog
An outlined modular cube with one orange corner block lifted above its place, on a pale blue background.
Published on
2026-10-08

12 distributed computing projects to build, from beginner to advanced

A distributed task queue is a useful first project: submit a job, send it to a worker, then kill that worker before it reports back. The next decision, whether and how to retry, introduces a problem that appears throughout distributed computing.

These 12 distributed computing projects progress from independent workers to replicated databases and failure testing. Each includes a small starting environment, the main engineering challenge, a success test, and an assessment of where Compute with Hivenet could help. A separate section covers volunteer computing for readers who want to contribute resources to scientific research.

Choose a project around one distributed problem

Pick the behavior you want to understand: recovering lost work, reconciling conflicting updates, agreeing on an operation order, or moving data while requests continue. Our guide to how distributed computing works explains the broader model. Here, the aim is to build something small enough that you can explain its failures.

Start with separate processes on one computer. They can exchange messages, time out, and crash independently as processes, which is enough to learn many protocol behaviors. They still share a host, network interface, and other resources. A local experiment does not establish resilience to independent machine failures.

The environments below are suggested starting configurations, not hardware requirements or production sizing advice. Use small inputs first. CPU resources are enough for these exercises unless you deliberately add a GPU workload. The Hivenet fit labels are editorial assessments: Good means a reasonable option for a self-managed remote experiment; Conditional means the design needs specific checks; Unnecessary means cloud infrastructure adds little to the starting exercise.

Beginner distributed computing projects

1. Distributed task queue

Difficulty: Beginner. Build a submission API, a coordinator that records job status, and workers that claim tasks and return results. Give each job an identifier. Add a time limit for a worker's claim and a retry policy for unfinished work.

  • Concepts: Work distribution, timeouts, idempotency, retries, and backpressure when jobs arrive faster than workers can finish them.
  • Starting environment: One coordinator and two worker processes on a laptop; later place workers on separate machines.
  • Main challenge: A missing response does not tell you whether a worker failed before or after producing a result. Retrying may execute the job twice.
  • Success test: Stop a worker mid-job, then verify that the job reaches an explicit completed or failed state. Repeat after the result is written but before acknowledgment, and check that retrying does not duplicate the intended output.

Hivenet fit: Good. Independent CPU workers are a reasonable first remote deployment. You still implement job coordination and protect the communication between workers and the coordinator.

2. Mini MapReduce engine

Difficulty: Beginner to intermediate. Split a text collection into input partitions, run map tasks, group intermediate results by key, and run reduce tasks. Word counts or a small inverted index make outputs easy to inspect. MIT's MapReduce lab offers a structured starting point with a coordinator and workers.

  • Concepts: Data partitioning, scheduling, intermediate data movement, slow workers, and failed-task recovery.
  • Starting environment: One coordinator, two workers, and a small local dataset. Decide how workers will retrieve intermediate files before moving to separate hosts.
  • Main challenge: A task can fail after writing only part of its output. Downstream work must not consume that partial result as complete.
  • Success test: Compare output with a single-process reference calculation. Interrupt a mapper and a reducer in separate runs; both runs should eventually produce the same correct output.

Hivenet fit: Good. Remote workers can process independent chunks, but you must arrange data access and task coordination. Deploying instances does not create a managed MapReduce service.

Intermediate distributed systems projects

3. Distributed web crawler

Difficulty: Intermediate. Build a shared URL frontier, several fetch workers, a deduplication mechanism, and a result store. Start against a test website you control, with known links, redirects, duplicate URLs, and deliberately slow responses.

  • Concepts: Shared work queues, duplicate suppression, per-site rate limits, backpressure, and recovery.
  • Starting environment: Two crawler workers, one coordinator or shared queue, and one small test web server.
  • Main challenge: Multiple workers can discover the same URL at once. A per-worker delay also fails to enforce a shared limit across the whole crawler.
  • Success test: Discover the expected pages, keep aggregate requests within your chosen per-site limit, and recover queued work after a worker crash without an uncontrolled wave of repeated requests.

Hivenet fit: Good. The workers can run as self-managed services. Before crawling external sites, respect robots.txt, applicable site terms, and request limits; keep the experiment bounded.

4. Replicated key-value store

Difficulty: Intermediate. Implement GET and PUT across three replicas. Start with a clearly documented replication policy, such as one write leader with asynchronous followers. Record versions so that stale reads and recovery behavior are visible.

  • Concepts: Replication, consistency, stale reads, quorum policies, and repair after a node returns.
  • Starting environment: Three storage processes with separate data directories and a test client that records requests and responses.
  • Main challenge: Define which writes count as successful, what a read may return, and when the service must reject requests.
  • Success test: Disconnect a replica, issue reads and writes, reconnect it, and check the recorded history against your declared guarantees. Demonstrate any stale reads instead of hiding them.

Hivenet fit: Conditional. Verify connectivity, latency, and disk behavior before using remote nodes. If you add quorum reads and writes, remember that overlapping quorums alone do not establish linearizability; version handling, concurrent writes, and failure rules matter too.

5. Distributed file store

Difficulty: Intermediate. Split files into chunks, track their locations in a metadata service, and keep two copies of each chunk on different storage processes. Add checksums and a repair worker. Treat the result as a learning prototype.

  • Concepts: Data placement, metadata, replication, integrity checking, and repair.
  • Starting environment: One metadata process and three storage processes, each with its own directory. Use files small enough to compare byte for byte.
  • Main challenge: Metadata can claim that a chunk exists even when a write failed. The metadata service also needs its own recovery story.
  • Success test: Remove one storage process while a valid replica remains. Retrieve the original file, verify its checksum, and restore the intended replication level. Separately corrupt a chunk and ensure the system detects it.

Hivenet fit: Conditional. Confirm the storage lifecycle and keep independent copies of important test data. This exercise requires your own file service and does not imply a managed distributed filesystem.

6. Distributed rate limiter

Difficulty: Intermediate. Make two application instances share a request budget. Begin with an atomic counter in a central limiter, then compare it with partitioned budgets or counters that synchronize periodically.

  • Concepts: Atomic updates, coordination overhead, clock assumptions, approximate counting, and availability.
  • Starting environment: Two application processes, a limiter service, and a load generator with a known request schedule.
  • Main challenge: Decide what happens when the limiter cannot be reached: allow requests, reject them, or consume a bounded local allowance.
  • Success test: Measure accepted and rejected requests during normal operation and a limiter outage. Report any budget overshoot and recovery behavior against the policy you chose.

Hivenet fit: Good. Small remote services let you observe coordination costs across real connections. Keep the workload modest until the limiter's accounting is correct.

7. CRDT collaborative editor

Difficulty: Intermediate to advanced. Use a conflict-free replicated data type (CRDT) to build a shared text editor whose clients can edit while disconnected and merge updates after reconnecting. You can first integrate a library such as Yjs, then implement a small replicated data type separately to study its merge rules.

  • Concepts: Concurrent updates, merge semantics, eventual delivery, and convergence.
  • Starting environment: Two browser clients and a local synchronization service, with a way to delay and replay document updates.
  • Main challenge: Identical final state does not guarantee that the merged text expresses what either author intended. Define the editing semantics as well as the delivery assumptions.
  • Success test: Edit both copies while disconnected, then deliver all updates with delays and duplicates. Check that every client converges and that subsequent edits still synchronize.

Hivenet fit: Unnecessary for the local exercise. It becomes a reasonable option if you later need a remote synchronization service for clients on different networks. A GPU is not needed for text synchronization.

8. Peer-to-peer file sharing

Difficulty: Intermediate to advanced. Let peers advertise and download file chunks from one another. Use content identifiers and checksums, then handle peers arriving or leaving. An explicit peer list is enough for version one; a distributed lookup mechanism can come later.

  • Concepts: Peer discovery, content addressing, integrity, incomplete availability, and network reachability.
  • Starting environment: Three peer processes with known addresses and test files replicated across more than one peer.
  • Main challenge: Knowing that a peer has a chunk does not mean you can connect to it. Firewalls and NAT complicate experiments across separate networks.
  • Success test: Remove a source peer during transfer while another source remains. Finish the download, verify the file, and reject an intentionally corrupted chunk.

Hivenet fit: Conditional. Check inbound and outbound connections, selected protocols, and port exposure. Do not assume that a public application endpoint supports every peer-to-peer protocol.

Advanced distributed computing projects

9. Raft replicated state machine

Difficulty: Advanced. Implement elections, terms, replicated logs, commit rules, persistent state, and recovery. Apply committed commands to a simple deterministic state machine. MIT's Raft lab provides a structured exercise and tests.

  • Concepts: Consensus, ordering, safety, and progress when a majority can communicate under suitable timing conditions.
  • Starting environment: Three peers and a controllable message transport. Begin locally with simulated delays, dropped messages, crashes, and restarts.
  • Main challenge: An isolated old leader may still believe it leads. Correct commit rules must prevent conflicting committed histories.
  • Success test: Isolate the leader, let a communicating majority elect a replacement, and issue commands. Reconnect the old leader and verify agreement on committed entries. Without a majority, new commands must not be reported as committed.

Hivenet fit: Conditional. Remote runs can extend local testing once storage and networking are understood. One deployment's timing measurements do not establish general Raft performance.

10. Sharded key-value database

Difficulty: Advanced. Partition keys across storage groups, then add a configuration service and controlled shard migration. Establish replication within each group before attempting migration under load.

  • Concepts: Partitioning, replicated ownership, reconfiguration, routing, and data transfer.
  • Starting environment: Two logical shard groups with three replica processes each, plus a configuration service and client. These can initially run on one host.
  • Main challenge: During a move, old and new owners must agree about which group may accept writes and which data has transferred.
  • Success test: Move a shard while clients read and write. Check that acknowledged writes remain available after recovery and requests obey the chosen consistency model. Explicit temporary rejection or retry can be part of the design.

Hivenet fit: Conditional. Plan data persistence, inter-node communication, and recovery before placing groups on remote instances. More nodes do not automatically produce useful independent failure domains.

11. Distributed job scheduler

Difficulty: Advanced. Extend the task queue with worker registration, resource reports, priorities, and placement decisions. Add expiring task assignments, often called leases, and recovery when a worker stops reporting.

  • Concepts: Resource allocation, scheduling, stale state, fairness, leases, and duplicate execution.
  • Starting environment: One scheduler and three worker processes with deliberately different capacity limits.
  • Main challenge: A delayed heartbeat can make capacity information stale. Reassigning work may overlap with a worker that is still executing its old assignment.
  • Success test: Submit mixed job sizes, verify the chosen capacity rules, and interrupt a worker. Account for every job and make repeated execution safe for its outputs.

Hivenet fit: Good for asynchronous workers. Treat this as your own scheduler experiment. Tightly coupled jobs need additional network and coordination checks; do not assume a managed Kubernetes or Slurm cluster.

12. Fault-injection test harness

Difficulty: Advanced. Build a controller that runs another project from this list, injects selected failures, and records the operation history. Include process crashes, delayed responses, duplicate requests, and interrupted communication.

  • Concepts: Failure models, observability, reproducible experiments, safety checks, and recovery testing.
  • Starting environment: One controller, a small target deployment you own, and application-level hooks or a test proxy for manipulating messages.
  • Main challenge: A fixed random seed helps reproduce decisions, but it does not guarantee identical scheduling on a real network. Record the actual event order and resulting history.
  • Success test: Expose a known bug in the target, retain the failing scenario, fix the bug, and show the same check passing. State which failures remain untested.

Hivenet fit: Conditional. Keep injected failures within resources you control and verify which controls your instance permits. Reliability testing does not by itself establish security or prove that every failure is handled.

Volunteer computing projects to join

Volunteer computing lets you run research tasks on your own hardware. You learn about resource limits, work queues, and task deadlines, but you do not implement the project's internal protocols simply by running its client.

The BOINC project directory lists research projects and supported platforms. Examples include Einstein@home in astrophysics, PrimeGrid in mathematics, and Rosetta@home in biology. A directory listing does not guarantee that suitable work is available today; check each project's announcements, requirements, and server status. Science United offers another way to participate through BOINC by choosing scientific areas to support.

Folding@home uses its own client, separately from BOINC, for simulations of protein motion. LHC@home, listed in the BOINC directory, supports CERN physics research. It should not be confused with the Worldwide LHC Computing Grid used by participating research institutions.

Check current participation instructions even for familiar names. SETI@home's website, checked in October 2026, states that the project is in hibernation and no longer distributing tasks, while analysis of existing data continues.

Start with hardware you already own and set limits appropriate to its temperature, power use, and your other work. Allow for electricity costs and task deadlines. Renting paid cloud instances to donate compute is a separate spending decision, not a requirement for volunteering.

When to use Compute with Hivenet

Move beyond your laptop when separate hosts answer a specific question: whether worker recovery survives a lost connection, how much remote coordination adds to latency, or how clients behave from different networks.

The Compute with Hivenet quickstart covers containers and virtual machines, vCPU-only and GPU options, and connection settings. Check available resources and current prices in the console. Choose CPU resources for these projects unless the work itself needs a GPU.

Before deploying stateful or peer-to-peer systems, test node-to-node reachability, required ports, storage behavior, and latency. Keep recovery copies outside the experiment: terminating an instance deletes its local data. You manage the application, coordination, access controls, and recovery; do not assume managed cluster software or a private high-speed network.

Keep experiment records alongside the code: topology, configuration, input data, failure schedule, observed behavior, and remaining limits. Our distributed systems management guide explains the operational responsibilities that grow as these prototypes become services.

Pick a first milestone you can verify

For a first project, complete the task queue's failure-and-retry test before adding a dashboard. For stateful systems, write down the read and write guarantees before adding replicas. For advanced work, build a controllable test environment before interpreting performance results.

A useful finished project has a repeatable setup, a small correct baseline, a documented failure case, and evidence that its recovery behavior matches its promises. Choose the smallest exercise that exposes the distributed problem you want to understand.

Your next workload belongs on Hivenet.

Pick one AI, compute, or storage workload and see the difference for yourself. Spin it up in minutes, or let our team map your fastest path to production.