Conversation
Signed-off-by: Juan Cruz Viotti <jv@jviotti.com>
There was a problem hiding this comment.
1 issue found across 33 files
You’re at about 91% of the monthly reviewed-line limit. You may want to disable incremental reviews to conserve quota. Reviews will continue until that limit is exceeded. If you need help avoiding interruptions, please contact contact@cubic.dev.
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="test/rdf/fail_jsonl_resolution_error.clitest">
<violation number="1" location="test/rdf/fail_jsonl_resolution_error.clitest:21">
P3: The `// Schema input error` comments mislabel this test: the schema itself is valid, and the failure is a JSON-LD resolution error (`x-jsonld-language` on the dataset entry), the same class the single-instance test `fail_resolution_invalid_language.clitest` covers. The label is copied from that test, which uses the same wording for the same error class, but here it will mislead readers skimming for schema-input failure cases (compare `fail_schema_not_a_schema`, `fail_schema_invalid_json`). Relabel both occurrences, e.g. `// Dataset resolution error`, to describe the scenario being asserted.</violation>
</file>
Reply with feedback, questions, or to request a fix.
Re-trigger cubic
| { "name": "Ada" } | ||
| EOF | ||
|
|
||
| // Schema input error |
There was a problem hiding this comment.
P3: The // Schema input error comments mislabel this test: the schema itself is valid, and the failure is a JSON-LD resolution error (x-jsonld-language on the dataset entry), the same class the single-instance test fail_resolution_invalid_language.clitest covers. The label is copied from that test, which uses the same wording for the same error class, but here it will mislead readers skimming for schema-input failure cases (compare fail_schema_not_a_schema, fail_schema_invalid_json). Relabel both occurrences, e.g. // Dataset resolution error, to describe the scenario being asserted.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At test/rdf/fail_jsonl_resolution_error.clitest, line 21:
<comment>The `// Schema input error` comments mislabel this test: the schema itself is valid, and the failure is a JSON-LD resolution error (`x-jsonld-language` on the dataset entry), the same class the single-instance test `fail_resolution_invalid_language.clitest` covers. The label is copied from that test, which uses the same wording for the same error class, but here it will mislead readers skimming for schema-input failure cases (compare `fail_schema_not_a_schema`, `fail_schema_invalid_json`). Relabel both occurrences, e.g. `// Dataset resolution error`, to describe the scenario being asserted.</comment>
<file context>
@@ -0,0 +1,57 @@
+{ "name": "Ada" }
+EOF
+
+// Schema input error
+RUN rdf schema.json instances.jsonl STDIN /dev/null IN . INTO result_0.txt EXPECTING 4
+
</file context>
🤖 Augment PR SummarySummary:
🤖 Was this summary useful? React with 👍 or 👎 |
|
|
||
| The instance may hold more than one document, as a [JSON Lines | ||
| (JSONL)](https://jsonlines.org) dataset, a GZIP-compressed JSONL dataset, or a | ||
| multi-document YAML file. Standard input is read the same way. Every document |
There was a problem hiding this comment.
docs/rdf.markdown:55 — The preceding list includes GZIP-compressed JSONL, but standard input is passed to read_stdin_documents, which only tries raw JSON, JSONL, and YAML and never enables GZIP mode. Consequently, piping a .jsonl.gz file to rdf ... - fails to parse despite the statement that standard input is read the same way.
Severity: medium
🤖 Was this useful? React with 👍 or 👎, or 🚀 if it prevented an incident/outage.
| on every line, which is what keeps each line a self-contained JSON-LD | ||
| document. | ||
|
|
||
| Documents are written as they pass, and the command stops at the first one |
There was a problem hiding this comment.
docs/rdf.markdown:74 — This is not true for a malformed later entry: read_instances fully parses and materializes for_each_json before the promotion loop begins. A valid first JSONL document followed by malformed JSON therefore produces no first-document output and reports the parse error rather than writing documents as they pass.
Severity: medium
🤖 Was this useful? React with 👍 or 👎, or 🚀 if it prevented an incident/outage.
There was a problem hiding this comment.
1 issue found across 5 files (changes from recent commits).
You’re at about 92% of the monthly reviewed-line limit. You may want to disable incremental reviews to conserve quota. Reviews will continue until that limit is exceeded. If you need help avoiding interruptions, please contact contact@cubic.dev.
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="docs/rdf.markdown">
<violation number="1">
P2: The removed paragraph documented behavior that still holds exactly: `read_instances` parses the whole input up front via `for_each_json` (see the TODO at src/command_rdf.cc:58-64), then the loop at src/command_rdf.cc:372-378 writes each valid entry and stops the run at the first entry that fails to validate or promote. Nothing left in this file says that a run exiting non-zero may already have printed the documents before the failure. Restore an equivalent note, otherwise readers will assume a failed dataset run produced no output or processed the whole dataset.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| @@ -2,7 +2,7 @@ Linked Data (RDF) | |||
| ================= | |||
There was a problem hiding this comment.
P2: The removed paragraph documented behavior that still holds exactly: read_instances parses the whole input up front via for_each_json (see the TODO at src/command_rdf.cc:58-64), then the loop at src/command_rdf.cc:372-378 writes each valid entry and stops the run at the first entry that fails to validate or promote. Nothing left in this file says that a run exiting non-zero may already have printed the documents before the failure. Restore an equivalent note, otherwise readers will assume a failed dataset run produced no output or processed the whole dataset.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At docs/rdf.markdown, line 75:
<comment>The removed paragraph documented behavior that still holds exactly: `read_instances` parses the whole input up front via `for_each_json` (see the TODO at src/command_rdf.cc:58-64), then the loop at src/command_rdf.cc:372-378 writes each valid entry and stops the run at the first entry that fails to validate or promote. Nothing left in this file says that a run exiting non-zero may already have printed the documents before the failure. Restore an equivalent note, otherwise readers will assume a failed dataset run produced no output or processed the whole dataset.</comment>
<file context>
@@ -72,13 +72,6 @@ a line rather than across the dataset, and `--compact/-c` repeats the context
-documents that preceded the failure, so check the exit code rather than the
-presence of output.
-
> [!NOTE]
> Annotation collection is a JSON Schema 2019-09 and 2020-12 feature, so this
> command requires the schema to have one of those dialects as its base
</file context>
Signed-off-by: Juan Cruz Viotti <jv@jviotti.com>
There was a problem hiding this comment.
1 issue found across 9 files (changes from recent commits).
You’re at about 92% of the monthly reviewed-line limit. You may want to disable incremental reviews to conserve quota. Reviews will continue until that limit is exceeded. If you need help avoiding interruptions, please contact contact@cubic.dev.
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="src/input.h">
<violation number="1" location="src/input.h:762">
P2: Optional callers do not convert an empty JSONL dataset into a missing-input error, so this condition suppresses the only diagnostic for commands such as `fmt empty.jsonl` and `lint empty.jsonl`. Preserve the warning for optional runs and suppress it only when a non-empty requirement already reports the empty result as missing input.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Re-trigger cubic
| std::size_t empty_datasets{0}; | ||
| auto result{for_each_json(arguments, options, requirement, formatting, | ||
| empty_datasets)}; | ||
| if (!result.empty()) { |
There was a problem hiding this comment.
P2: Optional callers do not convert an empty JSONL dataset into a missing-input error, so this condition suppresses the only diagnostic for commands such as fmt empty.jsonl and lint empty.jsonl. Preserve the warning for optional runs and suppress it only when a non-empty requirement already reports the empty result as missing input.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. At src/input.h, line 762:
<comment>Optional callers do not convert an empty JSONL dataset into a missing-input error, so this condition suppresses the only diagnostic for commands such as `fmt empty.jsonl` and `lint empty.jsonl`. Preserve the warning for optional runs and suppress it only when a non-empty requirement already reports the empty result as missing input.</comment>
<file context>
@@ -738,6 +742,30 @@ inline auto for_each_json(const std::vector<std::string_view> &arguments,
+ std::size_t empty_datasets{0};
+ auto result{for_each_json(arguments, options, requirement, formatting,
+ empty_datasets)};
+ if (!result.empty()) {
+ report_empty_datasets(empty_datasets);
+ }
</file context>
| if (!result.empty()) { | |
| if (requirement == InputRequirement::Optional || !result.empty()) { |
Signed-off-by: Juan Cruz Viotti jv@jviotti.com