Repository navigation
CBOR parser doesn't skip tags #1968
Description
Activity
Can you provide the error message?
- addedaspect: binary formatsBSON, CBOR, MessagePack, UBJSONBSON, CBOR, MessagePack, UBJSON
on Mar 5, 2020 It's this one here:
invalid byte: 0xD9.I‘v read the CBOR standard document and its implementation in the source code, it does not support data of major-type 6(Optional Tagging of Items). One of the reasons is that the other major-type data has satisfy various data types of JSON.
but user can parse to any kind of data, like0xd9d9f7...
55799 is one of tags, it serialized to 0xd9d9f7 by CBOR spec. As written in the specification:Tag 55799 is defined for this purpose. It does not impart any special semantics on the data item that follows; that is, the semantics of a data item tagged with tag 55799 is exactly identical to the semantics of the data item itself.
so can we add a case to skip it to avoid parsing error? @nlohmann
The documentation states that the library does not support tags at this point, so I would not speak of a bug. I am hesitant of skipping tags, because this would, to my understanding, silently change the type or intention of the encoding. Or, to put it differently, how would you propose mapping such read bytes to JSON values?
stale commented
on Apr 19, 2020 staleboton Apr 19, 2020 · Hidden as outdatedshow commentMore actions- addedstate: stalethe issue has not been updated in a while and will be closed automatically soon unless it is updatedthe issue has not been updated in a while and will be closed automatically soon unless it is updated
on Apr 19, 2020 I am hesitant of skipping tags, because this would, to my understanding, silently change the type or intention of the encoding. Or, to put it differently, how would you propose mapping such read bytes to JSON values?
Just ignore them. This is perfectly acceptable according to the standard:
A decoder that comes across a tag (Section 2.4) that it does not
recognize, such as a tag that was added to the IANA registry after
the decoder was deployed or a tag that the decoder chose not to
implement, might issue a warning, might stop processing altogether,
might handle the error and present the unknown tag value together
with the contained data item to the application (as is expected of
generic decoders), might ignore the tag and simply present the
contained data item only to the application, or take some other type
of action.- removedstate: stalethe issue has not been updated in a while and will be closed automatically soon unless it is updatedthe issue has not been updated in a while and will be closed automatically soon unless it is updated
on Apr 19, 2020 stale commented
on May 19, 2020 staleboton May 19, 2020 · Hidden as outdatedshow commentMore actions- addedstate: stalethe issue has not been updated in a while and will be closed automatically soon unless it is updatedthe issue has not been updated in a while and will be closed automatically soon unless it is updated
on May 19, 2020 It's a bit odd that you have a bot to close issues that nobody replies to after 1 month. I don't think it's going to magically fix itself!
- removedstate: stalethe issue has not been updated in a while and will be closed automatically soon unless it is updatedthe issue has not been updated in a while and will be closed automatically soon unless it is updated
on May 19, 2020 This is a one-man-show, and I do not want to be overwhelmed by inactive issues. Since I will never be able to work on all of them, I'd rather keep the active ones open.
3 remaining items
- removedstate: stalethe issue has not been updated in a while and will be closed automatically soon unless it is updatedthe issue has not been updated in a while and will be closed automatically soon unless it is updated
on Jul 10, 2020 Tags should be skipped to allow roundtripping after #2244 is merged.
Alright, I finally had time to work on this. My idea would be to pass the
from_cborfunction an enum parametertag_handlerthat would have two values:error: The current behavior of the library - in case a tag is read, aparse_errorexception is thrown. This is the default to avoid breaking existing client code.ignore: The tag is ignored and the following value is parsed as before. On case of tags0xD8-0xDB, the following 1-8 bytes are read, but also discarded. This would treat any tagged value as a "standard" CBOR value.
Also possible (but more difficult to implement would be):
binary: Tagged values are parsed to binary values. The subtype is the tag value and the following bytes are just copied to the binary value as is. I am not sure whether this is helpful.
Any comments on this?
- addedstate: please discussplease discuss the issue or vote for your favorite optionplease discuss the issue or vote for your favorite option
on Jul 12, 2020 - added a commit that references this issue
on Jul 12, 2020 - addedsolution: proposed fixa fix for the issue has been proposed and waits for confirmationa fix for the issue has been proposed and waits for confirmation
on Jul 15, 2020 Sounds sensible to me.
- removedstate: please discussplease discuss the issue or vote for your favorite optionplease discuss the issue or vote for your favorite option
on Jul 17, 2020
Using 456478b (3.7.3)
The following code should parse a CBOR value:
The value I'm parsing (
0xd9d9f7) is simply the optimal "magic number" tag for CBOR documents. From the specification:Byte 0xd9 should be fine because it is equal to
(6 << 5) | 25, in other words it has a major type of 6 (a tag), and lower 5 bits of 25, which for a tag means the actual tag value follows in auint16, so it should just skip the following 2 bytes.Some extra code needs to be added here. It doesn't understand tags at all.