Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
42 changes: 36 additions & 6 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,7 +14,7 @@ Package versions are derived from git tags (`v*`) via MinVer.
into the active browser profile, so they apply to every backend (static, JS, and headless). `--cookie`
now flows through the same path and reaches headless backends too.
- Adaptive per-host throttling (on by default): after repeated `429`/`503` responses a host's crawl
delay is raised — honouring the `Retry-After` header as a per-host grace — and eased back down on
delay is raised, honouring the `Retry-After` header as a per-host grace, and eased back down on
sustained success. Configurable via `ThrottleOptions` on `CrawlerOptions.Throttling`
(`Enabled`, `MaxDelaySeconds`) and the CLI flag `--adaptiveThrottle` (pass `false` to disable).
- Checkpoint/resume via `--checkpoint <file>`: the crawl frontier (discovered/processed/visited) is
Expand All @@ -24,11 +24,9 @@ Package versions are derived from git tags (`v*`) via MinVer.
`AbstractCrawler` gained an optional `ICheckpointStore` constructor parameter; autosave cadence is
configured via `CrawlerOptions.Checkpoint.Interval`.
- Per-URL reporting on `IScrapeResult` via a new `Reports` collection of `UrlReport` (in
`SimpleCrawler.Core.Models`, alongside the `CrawlOutcome` enum). Every fetched page — success or
failure — is reported with its status code, fetch/parse durations, content length/type, discovered
link count, index/follow flags, timestamp, outcome, and any error. `Urls` is unchanged (still the
indexable subset). Reports live in `CrawlState`, so they are checkpointed and restored on resume in
step with `Urls`.
`SimpleCrawler.Core.Models`, alongside the `CrawlOutcome` enum). Every fetched page is reported with
its status code, fetch/parse durations, content length/type, discovered link count, index/follow flags,
timestamp, outcome, and any error. `Urls` is unchanged (still the indexable subset).
- Optional `--report <file>` CLI flag that writes the per-URL report as JSON. The existing plain
URL-per-line output is unchanged.
- Opt-in WHATWG Streams surface for the JS backends via `JsRenderOptions.EnableStreams` (off by
Expand All @@ -39,6 +37,11 @@ Package versions are derived from git tags (`v*`) via MinVer.
bundle (e.g. Next.js App Router RSC) can otherwise tear down the server markup without rebuilding it
in this single-pass render, so the renderer captures a pre-script anchor baseline and restores it at
finalize when the live tree regresses below the shell's links.
- Surfaced exception diagnostics for the JS renderer. Catches now route through an unconditional
`__crawlerDiagnostic` channel at `Debug` level, so raising the log level to `Debug` turns
every silent settle into a named exception with a stack. The message is emitted before
the stack: Jint's `error.stack` is frames-only, so reporting the stack alone dropped
what identifies the failure.

### Changed

Expand All @@ -59,13 +62,40 @@ Package versions are derived from git tags (`v*`) via MinVer.
C#↔JS DOM bridge (cut in Phase 6), never invoked by any renderer path or test. The unreachable
`__crawlerReset` realm-reset machinery was also deleted from the JS DOM prelude; it existed only for
the since-removed Jint realm pool and is internal cleanup.
- The Jint `Map.keys()`/`values()` iterator compat shim and its shims prelude. Jint 4.11 fixed the
"Collection was modified" bug it patched (a bundle mutating a `Map` during iteration), so the
per-engine shim is no longer needed.

### Fixed

- robots-meta parsing started `index`/`follow` at `false`, so a lone `content="index"` (no explicit
`follow`) parsed as `follow=false` and the crawler dropped every link on the page. The spec defaults
are `index, follow` and directives only negate, so the flags now start `true`; lone and combined
directives keep their permissive defaults. Pinned by `IndexingHelperTests`.
- RSC-adjacent DOM gaps that crashed Next.js App Router (RSC) sites during hydration/commit. The shim
gaps the crashes traced to are closed: `Element.attributes` is now a live `NamedNodeMap`-backed
collection that shrinks as attributes are removed (React's singleton teardown loop relied on the
collection actually contracting), with real `removeAttributeNode`/`getAttributeNode`/`setAttributeNode`,
plus `Node.getRootNode`, `document.getElementsByName`, a global `reportError`, and `document.readyState`
with a `DOMContentLoaded` dispatch after bundle execution. App Router RSC sites now render fully on the
default V8 engine; on Jint they still fail with "Cannot convert undefined or null to object" from React's
Flight deserialization — an upstream engine bug tracked separately.
- Checkpoint/resume was only reachable from `AbstractCrawler`-based (static) crawlers: the concrete
AngleSharp, JS (Jint/V8), Playwright, and Puppeteer constructors didn't forward an `ICheckpointStore`
parameter, so DI could never supply one for those backends. All backend constructors now accept and
thread through an optional `ICheckpointStore`.
- `Ctrl+C` cancelled the crawl token but let .NET terminate the process immediately afterward, racing
the in-flight checkpoint save. The CLI now cancels on the first `Ctrl+C`, waits for in-flight requests
to finish and the checkpoint to persist, and logs that it's doing so; a second `Ctrl+C` exits immediately.
- JS DOM: `HTMLElement` had no Constraint Validation API and there was no global `FileList`, so frameworks
that call `setCustomValidity`/`checkValidity`/`reportValidity` on a form-control ref, or touch a file
input's `.files`, threw during hydration. Added a no-op-but-spec-shaped Constraint Validation API
(`willValidate`, `validity`, `checkValidity`, `reportValidity`, `setCustomValidity`) and a `FileList` global.
- Jint's `Function.prototype.toString()` defaulted to a hardcoded `"[native code]"` stub for every
ordinary script function (unlike V8/real browsers, which print real source), making bundle-authored
functions indistinguishable from the host DOM methods `browser/native.ts` deliberately marks as native
for jQuery/Sizzle's native-code sniff. Prepared scripts/modules are now tagged with their source text
at parse time so `Options.Host.FunctionToStringHandler` can return the real text for everything else.

## [2.0.0] - 2026-07-07

Expand Down
10 changes: 10 additions & 0 deletions docs/javascript-crawlers.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,10 +45,20 @@ Passed to `AddJintCrawler`/`AddV8Crawler`.
| ------ | ------- | ------ |
| `EnableFetch` | `false` | Enables real network `fetch`/`XHR` for runtime-loaded content/links to render. |
| `EnableIndexedDb` | `false` | Installs an in-memory `indexedDB`. Turn it on for SPA sites that rely on offline features. |
| `EnableStreams` | `false` | Installs a [WHATWG Streams](https://streams.spec.whatwg.org/) surface. Delivering spec-compliant reader/transform semantics, not incremental transport streaming. See [RSC and streaming bundles](#rsc-and-streaming-bundles). |
| `Viewport` | 1920×1080 | Window/screen reported to scripts. Set a mobile screen size to crawl the mobile layout on responsive sites. |
| `ScriptLogging` | `null` | `LogLevel` floor for forwarding page `console.*` to your logger. `Debug` to diagnose non-rendering pages. |
| `MaxTaskDrainIterations` | `1000` | Cap on microtask/chunk-load drain iterations before giving up on a page. |

## RSC and streaming bundles

Next.js App Router (React Server Components) sites render on the **V8** engine.

The same sites still fail on **Jint** with "Cannot convert undefined or null to object" from React's Flight deserialization, this is a bug in Jint [#2607](https://github.com/sebastienros/jint/pull/2607), use V8 (the default) for RSC sites while that is fixed.

`EnableStreams` delivers a buffered-complete body (the fetch already materialises the whole response), so consumers get spec-compliant reader/transform semantics, not chunks-over-time, and the baseline guard keeps the server markup intact if the
streaming path would otherwise tear it down. Sites that need real streaming should use the [Playwright](./configuration.md) or [Puppeteer](./configuration.md) headless backends.

## Engine reuse

The two engines handle per-page setup differently.
Expand Down
Loading
Loading