Our services

If you can think it, we can make it brainsoft.

Keeping two systems in sync without a nightly batch job

Written By: BrainSoft In Backend

A nightly batch job that copies data between two systems is easy to build and easy to trust, right up until it silently fails on a Tuesday and nobody notices until a customer complains that their order is missing from the second system. The job "worked" in the sense that it ran. It did not work in the sense that mattered.

The fix is not a fancier batch job. It is moving the integration closer to real time, making every write safe to repeat, and logging enough that a mismatch six months from now is a ten-minute investigation instead of a day of guessing.

Why nightly batch jobs go wrong

Batch jobs fail quietly because nothing is watching them closely enough to notice a partial failure. A job that processes ten thousand records and fails on record four thousand often reports success, because "the script finished" and "the script did what it was supposed to" get treated as the same thing. The gap between them is where the missing order lives.

Webhooks first, polling as a fallback

If the source system can call you when something changes, use that. A webhook that fires on every order, update or cancellation keeps both systems close to in sync and fails loudly — a missed webhook shows up as an error in your logs within seconds, not as a gap discovered days later. Keep a polling job as a safety net that reconciles anything the webhook missed, but treat it as a backstop, not the primary path.

Make every write idempotent

Any integration will eventually retry a request — a timeout, a redeploy mid-job, a webhook delivered twice. If replaying the same event twice creates a duplicate order or double-charges a customer, the integration is not safe to run in production. Give every event an ID, record which IDs you have already processed, and make the write a no-op the second time it arrives.

Log enough to debug a mismatch six months later

When a customer disputes an order total, you want to know: what payload arrived, when, what your system did with it, and whether it succeeded. Store the raw payload alongside the processed result, not just the outcome. Storage is cheap; reconstructing what happened from application logs after the fact usually is not possible at all.

This is the kind of work we do as part of enterprise integration projects — connecting existing systems without creating a second source of truth nobody trusts. If you are staring at a batch job you no longer trust, get in touch and we can take a look.

Frequently asked questions

Is a nightly batch job ever the right choice?

Yes, for data that genuinely does not need to be current within the day, such as historical reporting or analytics warehouses.

What is idempotency, in plain terms?

Running the same operation twice produces the same result as running it once, which stops a retried request from creating a duplicate record or a second charge.

How do we find out our two systems are already out of sync?

Run a reconciliation job that compares record counts and key totals between the two systems on a schedule, and alert on drift instead of waiting for a customer to report it.


#Backend