Nobody changed the field. The default did.
An update that never mentioned a field overwrote it anyway, and the check we had trusted since the last incident could not see it. Two days later a job that was never told where to start read seven weeks of history and reported a backlog of 106 items. Run with the real start time, it found 2. Both were defaults doing exactly what they were written to do.
We keep the opportunities we track in a database table and write to it through a small command-line tool. In August that tool deleted data. An update meant to add one note to one record removed ten of that record's twenty fields, then printed its normal success line and exited cleanly. We found out only because we read the record back afterward.
The update command was called upsert, a name that promises a merge. It was a replace. It built a new record from whatever the caller passed, plus three fields it had been told to protect, and wrote that over the stored one. Everything else the caller left out was gone. All twenty fields were restored straight away, though one of them could not be rebuilt from anything left on the record and had to be re-read from the original notice.
The fix was the textbook one. Read the stored record first, lay the caller's values over it, and write the result. A value the caller supplies wins. A value the caller leaves out keeps whatever was stored, and deleting a field became something you have to ask for by name. We tested it on the exact call that had done the damage, twenty fields before and twenty after, and later re-checked it on a live record, where an update passing three fields left all nineteen of its fields in place.
The field that changed without being mentioned
Last week we learned that the nineteen-field result had been true while the tool was overwriting data.
One field on each record decides what happens once a finished response is approved: it goes out by email, or it lands on a list to be submitted by hand through a portal. The tool gave that field a default of manual submission, and the comment beside it explained why that was the safe choice. A record marked manual is never sent automatically.
The comment said new records. The code said any call that left the field out.
The code also applied the default before the merge. By the time the merge asked whether the caller had supplied a value, the answer was always yes, so every update that did not mention the field wrote manual over whatever was stored. When we checked, three of the five responses queued to go out by email sat on records marked for submission by hand. All five were still waiting for approval, and the records were corrected the same day.
Why the check we trusted could not see it
After the August incident, our check for this tool was to read a record before and after a write and compare the set of fields. That check is exactly right for the failure it was built to catch. When a write deletes fields, the set shrinks and the difference is obvious.
A value that changes in place leaves the set alone. Same names, same count, different contents. The nineteen-field result was a true statement about the thing we were measuring and silent about the thing that was going wrong. A count of fields cannot see a changed value, and neither can a success message.
It was also the second time. Eleven days earlier the same shape had broken a field meant to be written once: the record of which piece of work first created each row. The comment above it said write-once held by construction, since the merge keeps anything the caller leaves out. It did not account for the environment. When a caller did not pass the field, the tool filled it from an environment variable, and in normal operation that variable is always set. A routine backfill rewrote the creator on two records and stamped one onto two older records that had never had one. Reading the value back after the write is what caught it. The write itself had printed success.
Both defects have the same anatomy. A merge keeps what a caller leaves out only if the omission survives until the merge runs. A default applied earlier fills the gap, and after that the merge cannot tell a value the caller chose from one the code chose for them.
What we changed
The repair was small. The tool now records whether it applied the default, and when it did, the merge keeps the stored value instead. A new record still gets the safe default, and a value the caller passes explicitly still wins. The creator field got the same rule: only the write that creates a record fills it, unless the caller sets it on purpose.
The test matters more than the fix. It replays the call that caused the reset against real records, without writing anything, and checks the value that would be stored. Before trusting it we ran it against the old code, which produced manual, so we know the test can tell the two versions apart. A test that passes on the broken code and the fixed code alike is not testing the fix. We also went through the rest of that write path looking for any other default applied before the merge, and found none left unguarded.
The same mistake in a job that writes nothing
Two days later the same shape showed up somewhere with no record to overwrite.
A scheduled job of ours reads a message feed from another system and reports what has arrived since its last run. The caller passes the time of that last run as a parameter. When the start time was first made a parameter, in August, the old hard-coded date stayed behind as a fallback for any call that did not pass one. The comment explained why: the fallback preserved the old behaviour, rather than silently reporting everything as new.
The wrapper that runs the job on its remote machine passes no arguments. A note in the file said exactly that, and showed the right way to call the job, twelve days before a scheduled run called it bare anyway. The job fell back to the August date, read a window of 1,262 hours, ran into its own time limit, and reported a seven-week backlog of 106 items needing action. Run again the next day with the real start time, it found 2.
The fallback was there to stop the job reporting everything as new. It reported seven weeks as new.
The note was accurate and it did not help, because a note only works on whoever reads it at the moment it matters. We removed the fallback instead. A call with no start time now stops with an error that prints the correct invocation. The test for the change fails three of its seven checks against the old script, and the normal path, run through the real wrapper before and after the change, produced byte-identical output. The fix changed what happens on the failure path and nothing else.
What a default is allowed to say
A few weeks ago we wrote about a record with a missing field, which a watcher skipped without a word. These cases are the mirror image: values that were not missing at all, overwritten or replaced because nobody mentioned them.
A default is a value written in advance for a caller who is not there. It is honest only where the thing it stands in for does not exist. On a brand-new record, “no submission method yet” is simply a fact, and a default fills it truthfully. On an existing record, “this update did not mention it” is a different fact, and a default there is an overwrite. For the scheduled job, a missing start time did not mean there had been no previous run. The real time was on record, and the default replaced it with a guess seven weeks old.
When no guess is safe, the only safe default is to stop and say so.
The version that works is not complicated. When we added a creation timestamp to these records in August, the obvious implementation would have been rewritten on every update: a field that looked like an intake time while behaving exactly like the last-updated time. Instead the write reads the stored record and fills the timestamp only when there is none. Two live writes four seconds apart confirmed it. The creation time stayed where the first write put it, and the update time moved.
We also left the older records without one. Filling them in from the last-updated time would have produced a creation history that was confident and wrong, and an invented date looks exactly like a real one. A visible gap in the data is a better record than a plausible guess.
For anyone who owns a write path
The questions are short. Where does your code supply a value nobody passed? For each one, does it apply only when a record is created, or on every write? By the time your merge runs, can it still tell a value the caller left out from one the caller set? When a required input is missing and no guess is safe, does the job stop, or does it pick a date?
And when you verify a write, what do you compare? A count of fields catches deletions. It will not catch a value that changed in place, and neither will the write's own success message. Read back the value you care about.
A default is a small decision made in advance for someone who is not in the room. Make sure it only speaks when there is genuinely nothing else to say.
The engineering behind the story.
We build data pipelines where a write does what it says. An update that leaves a field out keeps it. A default applies only where a value is genuinely absent, and a missing input stops the job instead of being guessed. When a write reports success, the proof is a read-back of the value that matters, not the report.