Why a NetSuite approval rolled itself back under load, and how a sweep fixed it
- Published on
- -7 mins read
- Authors
- Name
- Andy Nur (Andy)
- @andynur
Some bugs only appear when the account is busy. This one passed every Sandbox test with ten records, then showed up the first time finance bulk approved a few hundred expense reports in Production: some reports were no longer approved, even though the bulk approval batch had reported success.
In this article, I'll walk through the root cause and the fix. Client details are anonymized and the code is simplified.
We'll cover:
- How the approval queued a Map/Reduce through a workflow action
- How
NO_DEPLOYMENTS_AVAILABLEturned intoRCRD_HAS_BEEN_CHANGED - Why the fix is "don't save", plus a scheduled sweep
- Making the sweep safe to overlap with the normal path
Prerequisites
You should know SuiteFlow workflows, workflow action scripts, Map/Reduce scripts, and N/task. Background on the bulk approval tool helps, but isn't required.
The setup
Some expense reports have lines charged to another subsidiary. When such a report is approved, an intercompany journal must be created. Creating it takes real work, so the approval workflow doesn't do it inline. A workflow action script queues a Map/Reduce for that one report, and returns:
// Workflow action script, runs inside the approval transitionfunction onAction(context) { const er = context.newRecord try { task .create({ taskType: task.TaskType.MAP_REDUCE, scriptId: 'customscript_create_ic_journal_mr', params: { custscript_expense_report_id: er.id }, }) .submit() } catch (e) { // Original version: write the error to the report so someone sees it record.submitFields({ type: record.Type.EXPENSE_REPORT, id: er.id, values: { custbody_ic_journal_error: e.message }, }) }}With one approval at a time, this works.
What happened under load
A Map/Reduce script can only run as many tasks at once as it has free deployments. During a bulk approval, hundreds of approvals each queue a task within minutes. Once every deployment is busy, task.submit() throws NO_DEPLOYMENTS_AVAILABLE.
That alone would only mean a missing journal. The real damage came from the catch block. Step through it:
How a busy queue undid an approval
Step 1 / 61. Approve ER-1042
One of hundreds of approvals in the batch. The approval transition starts and the report is being saved.
In short, an optional side effect (the intercompany journal) could undo the main action (the approval). It only happened when the queue was full, which is why small tests never caught it.
The fix, part 1: never save inside the approval
The first change is small. For expense reports, the workflow action no longer writes anything back when queuing fails. A busy queue is expected during bulk runs, so it's logged as an audit entry, not an error:
} catch (e) { if (recordType === record.Type.EXPENSE_REPORT) { if (e.name === 'NO_DEPLOYMENTS_AVAILABLE') { log.audit('IC journal deferred', `ER ${er.id}: queue busy, sweep will create it`) return } log.error('IC journal not queued', e) return } // other record types keep the original behavior}The approval now always commits. The report may be approved without its journal for a while, which is what part 2 handles.
Rule of thumb
Code that runs inside a record's save, such as a workflow action or a user event, must not save that same record again. Log, queue, or defer, but don't write back.
The fix, part 2: a sweep
The same Map/Reduce now has a second mode. When it runs with no record parameter, it sweeps: it finds approved expense reports that should have a journal but don't, and creates them. A scheduled deployment runs it regularly.
function getInputData() { const id = runtime.getCurrentScript().getParameter({ name: 'custscript_expense_report_id' }) if (!id) return sweepPendingReports()
const er = record.load({ type: record.Type.EXPENSE_REPORT, id }) // The sweep may have created it while this task waited in the queue if (er.getValue('custbody_ic_journal_created')) return [] return buildJournalChunks(er)}Making the sweep safe
A sweep that runs next to the normal path can easily create duplicates. These rules keep it safe:
- Minimum age. Only reports approved at least 10 minutes ago. A task that's still in the queue gets time to finish first.
- Lookback and batch size. At most 7 days back and 25 reports per run, so one run stays well inside governance.
- Skip known errors. Reports with a recorded journal error are left for the manual Retry button. The sweep only handles "never queued".
- Check for an existing journal. Before building anything, exclude reports that already have a journal pointing back to them.
- Re-check in the normal path. The per-report task checks the "journal created" flag at start, as shown above.
- One sweep deployment. Only one scheduled deployment exists, so two sweeps never overlap.
function findPendingReportIds(limit) { const ids = [] search .create({ type: search.Type.EXPENSE_REPORT, filters: [ ['mainline', 'is', 'F'], 'AND', ['approvalstatus', 'anyof', '2'], 'AND', ['custbody_ic_journal_created', 'is', 'F'], 'AND', ['custbody_ic_journal_error', 'isempty', ''], 'AND', ['custbody_approved_date', 'onorafter', 'daysago7'], 'AND', ['custbody_approved_date', 'onorbefore', 'minutesago10'], 'AND', ['custcol_intercompany_path', 'noneof', '@NONE@'], ], columns: [search.createColumn({ name: 'custbody_approved_date', sort: search.Sort.ASC })], }) .run() .each((r) => { if (ids.indexOf(r.id) === -1) ids.push(r.id) return ids.length < limit }) return ids}The search runs on line level (mainline is false) because the intercompany flag lives on lines, so the same report can appear more than once. The indexOf check removes duplicates.
The sweep filters and the skip paths have Jest tests.
Conclusion
The bug wasn't in any single line of code. Each piece was reasonable on its own: queue heavy work, record errors on the record. Together, under load, they turned a full queue into lost approvals.
The fix follows a pattern that works well in NetSuite:
- Keep the main action (the approval) independent of its side effects
- Make side effects eventually consistent
- Add a sweep that finds and repairs what was missed, with guards against duplicates
If you queue tasks from workflow actions or user events, test with more records than you have deployments.