Togoder security

Comparison

npm audit vs source code scanning

npm audit answers whether anyone has reported a problem with your dependency versions. Source code scanning answers what the code in those versions actually does. They catch different things, and malware mostly lives in the gap between them.

Tooling6 min readUpdated By Togoder Security

Key takeaways

  • npm audit, Dependabot alerts and osv-scanner match your dependency versions against databases of already-reported vulnerabilities and malicious packages.
  • Advisory scanners cannot detect a malicious version until someone has found it and published an advisory, which can take hours to weeks.
  • Source code scanning analyzes what a package's code does, such as install scripts, network calls, credential reads and obfuscation, so it can flag unreported malware.
  • Behavioral and AI-based scanners produce more false positives than advisory matching and do not replace vulnerability (CVE) management.
  • The practical answer is to use both: advisory scanning for known CVEs and source analysis for new malicious code.

"We run npm audit" is the most common answer to "how do you check your dependencies?", and it is a good start. But npm audit was designed for a specific problem: known vulnerabilities in legitimate packages. Malicious packages are a different problem. This guide explains the two categories of tooling, what each one catches and misses, and how to combine them.

Two different questions

Advisory scanning asks: is package@version on a list of known-bad versions? The list is curated by humans and security vendors. A match is high-confidence; a miss means only "nobody has reported this yet".

Source or behavioral scanning asks: what does this code do when installed or imported? Does it read ~/.npmrc, post environment variables to a server, decode and eval a blob, spawn a shell? It works on code nobody has looked at before, at the cost of needing judgment about intent.

Advisory and CVE scanners

npm audit

Built into npm. It sends your dependency tree to the registry and matches it against the GitHub Advisory Database, which includes both CVEs and malware advisories (GHSA entries for malicious packages).

npm audit
npm audit --omit=dev
npm audit --audit-level=high   # non-zero exit only for high/critical
npm audit --json

Strengths: zero setup, fast, authoritative for known issues, gives fix versions. Weaknesses: noisy for vulnerabilities in dev-only or unreachable code, and completely blind to anything not yet in the database.

Dependabot, Renovate and similar

Dependabot alerts use the same GitHub Advisory Database and open pull requests to upgrade. Renovate focuses on keeping dependencies current and can surface vulnerability information. They are about remediation workflow more than detection. An important side effect: automated update PRs pull in brand-new releases quickly, which is exactly when a hijacked version is most dangerous, so configure a minimum release age where possible.

OSV and osv-scanner

OSV is an open, cross-ecosystem vulnerability database that aggregates GitHub advisories, ecosystem databases and others, including the OpenSSF malicious-packages feed. osv-scanner reads lockfiles for npm, PyPI, Go, crates.io, RubyGems, Packagist and more:

osv-scanner --lockfile=package-lock.json
osv-scanner -r .

Same model as npm audit, broader ecosystem coverage, and easy to run in CI. Commercial software composition analysis (SCA) products such as Snyk also sit mostly in this category, adding their own curated databases, reachability analysis and license checks.

Source and behavioral scanners

These tools download the package contents and analyze them. Approaches vary:

  • Static rules and heuristics. Flags for install scripts, network access, filesystem access, shell execution, obfuscated code, newly added maintainers, and so on. Socket.dev is the best-known product in this category for npm and other ecosystems, surfacing such signals on pull requests. Open-source options like OpenSSF Package Analysis run packages in a sandbox and record behavior.
  • Dynamic analysis. Install the package in an instrumented sandbox and watch syscalls and network traffic. Catches what code actually does at install, but can be evaded by payloads that check their environment or wait.
  • LLM source review. Have a language model read each file and judge whether the behavior is consistent with the package's purpose. This is what Togoder Security does: a cheap triage model clears obviously benign files, then a full model reviews the rest for install-time payloads, credential and wallet theft, exfiltration, obfuscation, backdoors and miners. Results are cached by file hash, so shared files across projects are reviewed once.

Strengths: can flag a malicious version minutes after it is published, before any advisory exists. Weaknesses: false positives on legitimate code that looks suspicious (installers that download binaries, telemetry, CLIs that spawn processes), possible false negatives on well-hidden logic, and they do not track ordinary CVEs like a prototype pollution bug in a legitimate library.

Side-by-side comparison

Advisory scanning (npm audit, Dependabot, osv-scanner, SCA)Source / behavioral scanning (Socket-style, sandboxing, AI review)
InputPackage names and versionsPackage file contents
Detects known CVEsYes, primary purposeGenerally no
Detects reported malwareYes, once an advisory existsYes
Detects unreported malwareNoYes, with varying accuracy
Time to detection for new malwareHours to weeks after publicationAs soon as the version is scanned
False positivesLow on identity, high on relevanceHigher; needs human triage
Fix guidanceUpgrade to version XRemove, pin, or replace the package
Cost to runUsually freeFree tiers to paid; compute-heavy

Why the detection gap matters

Look at how real incidents unfolded. When a maintainer's account was phished in September 2025 and malicious versions of chalk and debug were published, the community noticed and the versions were removed within a few hours. For anyone who ran a fresh install during that window, npm audit would have reported nothing, because the advisories did not exist yet. The Shai-Hulud worm made the window even more dangerous: each infected developer could publish new trojanized packages automatically, each one starting a fresh window.

The flip side is that most installs happen long after a version is published. If you install a malicious version that was reported last month, advisory scanning catches it with certainty and source scanning is redundant. That is why neither category is enough alone.

How to combine them

  1. Every PR and nightly: run npm audit --audit-level=high or osv-scanner against the lockfile. Fail builds on known malware advisories regardless of severity filters.
  2. When the lockfile changes: run a source-level scan of new and changed packages. Only the diff needs review, which keeps cost and noise low. See scanning your lockfile in CI.
  3. Before adopting any new version: wait a few days. Both categories get better with time: advisories get filed, and behavioral scanners have had a chance to look.
  4. Triage source findings by hand. A flag means "read this file". Our reports point to the exact file and describe the behavior; the methodology page explains how verdicts are reached.

You can see what source-level findings look like on public reports such as express or react, and on the malicious packages tracker.

Questions to ask any dependency scanner

Whichever tools you pick, these questions separate useful coverage from a checkbox:

  • What exactly does it read? Version strings only, the package manifest, or every file in the published tarball? Malware hidden in a minified bundle or a file referenced only from a postinstall script is invisible to manifest-level checks.
  • Does it analyze the published artifact or the repository? event-stream's payload existed only in the published flatmap-stream build, not in readable source on GitHub. Only artifact-level analysis would have seen it.
  • Does it cover transitive dependencies? Most incidents, including coa, rc and node-ipc, reached victims as dependencies of dependencies. A scanner that only checks direct dependencies misses most of the tree.
  • Which ecosystems? If your repo also has a requirements.txt, Cargo.lock or go.sum, an npm-only tool leaves gaps. Python .pth files and Rust build.rs scripts are install-time vectors just like npm lifecycle scripts.
  • How are findings explained? A useful finding names the file, the behavior and why it is suspicious, so a reviewer can confirm or dismiss it in minutes. An unexplained risk score just gets ignored.
  • What happens to your data? Lockfiles reveal your dependency list and sometimes internal package names. Check whether the service stores them and who can see results.

Bottom line

npm audit is necessary and not sufficient. It is the right tool for known vulnerabilities and known malware, and it costs nothing. Source code scanning is the right tool for the first hours and days of a malicious release, when no database knows about it yet. Use both, and treat any automated verdict, from either category, as input to a human decision.

Frequently asked questions

Is npm audit enough to protect against malicious packages?

No. npm audit only knows about packages with published advisories, so newly published malware passes until it is reported. It should be combined with source-level scanning.

What is the difference between osv-scanner and npm audit?

Both match dependency versions against vulnerability databases. osv-scanner uses the cross-ecosystem OSV database and supports many lockfile formats, while npm audit uses the GitHub Advisory Database and only covers npm.

Does source code scanning find CVEs?

Generally not. Behavioral and AI source scanners look for malicious intent such as credential theft or exfiltration, not ordinary bugs in legitimate code, so you still need an advisory scanner for CVEs.

Why do source code scanners have more false positives?

Legitimate packages also download binaries, spawn processes and make network requests. A scanner has to judge intent from code, which is harder than matching a version string against a list.