Our services
If you can think it, we can make it brainsoft.
If you can think it, we can make it brainsoft.
Written By: BrainSoft In Backend
A nightly batch job that copies data between two systems is easy to build and easy to trust, right up until it silently fails on a Tuesday and nobody notices until a customer complains that their order is missing from the second system. The job "worked" in the sense that it ran. It did not work in the sense that mattered.
The fix is not a fancier batch job. It is moving the integration closer to real time, making every write safe to repeat, and logging enough that a mismatch six months from now is a ten-minute investigation instead of a day of guessing.
Batch jobs fail quietly because nothing is watching them closely enough to notice a partial failure. A job that processes ten thousand records and fails on record four thousand often reports success, because "the script finished" and "the script did what it was supposed to" get treated as the same thing. The gap between them is where the missing order lives.
If the source system can call you when something changes, use that. A webhook that fires on every order, update or cancellation keeps both systems close to in sync and fails loudly — a missed webhook shows up as an error in your logs within seconds, not as a gap discovered days later. Keep a polling job as a safety net that reconciles anything the webhook missed, but treat it as a backstop, not the primary path.
Any integration will eventually retry a request — a timeout, a redeploy mid-job, a webhook delivered twice. If replaying the same event twice creates a duplicate order or double-charges a customer, the integration is not safe to run in production. Give every event an ID, record which IDs you have already processed, and make the write a no-op the second time it arrives.
When a customer disputes an order total, you want to know: what payload arrived, when, what your system did with it, and whether it succeeded. Store the raw payload alongside the processed result, not just the outcome. Storage is cheap; reconstructing what happened from application logs after the fact usually is not possible at all.
This is the kind of work we do as part of enterprise integration projects — connecting existing systems without creating a second source of truth nobody trusts. If you are staring at a batch job you no longer trust, get in touch and we can take a look.
Yes, for data that genuinely does not need to be current within the day, such as historical reporting or analytics warehouses.
Running the same operation twice produces the same result as running it once, which stops a retried request from creating a duplicate record or a second charge.
Run a reconciliation job that compares record counts and key totals between the two systems on a schedule, and alert on drift instead of waiting for a customer to report it.