↵ select ↓ ↑ navigate esc close

[RSS Club] Sorry for breaking your feed readers!

Terence Eden’s Blog ·

You're part of the Groovy Gang because you're a member of RSS Club! These posts are only available on my RSS and Atom feed. This post is not available in the shops, on the web, via FTP, or anywhere else.

So, yeah, sorry! My last post apparently broke some people's RSS readers. I had a code sample which said <marquee> - despite being properly escaped, some feed readers double-decoded it and turned it into a literal marquee element!

</video?

With thanks to Neil and CaféHaine for the videos.

I got several reports that people's readers started scrolling like that and they'd raised issues with their feed reader. Ooops! Sorry!

That said, as far as I can tell, the feed is escaped correctly and shouldn't cause problems.

Here's the code (I've added in some spaces to ensure it doesn't cause any issues):

1<content type="html">
2   <![CDATA[< p> Lorem ipsum <code>& lt;marquee&gt;</code> dolor sed.</p>

So what's going on? The feed is generated by the latest version of WordPress which uses CDATA to wrap HTML in a feed.

There is a 17 year old discussion about whether this is conformant on the WordPress issue tracker with the conclusion that it isn't incorrect and seems to work fine.

Is it OK? Is my feed broken or are a bunch of readers non-compliant? Let's go back to basics. The Atom spec says

If the value of "type" is "html", the content of atom:content MUST NOT contain child elements and SHOULD be suitable for handling as HTML. The HTML markup MUST be escaped; for example, "<br>" as "&lt;br>".

RFC 4287: The Atom Syndication Format

Hmmmm. That would indicate that ampersand-l-t-semicolon should be interpreted as a less-than sign.

However, the whole thing is wrapped in <![CDATA[ which according to the XML spec means:

CDATA sections may occur anywhere character data may occur; they are used to escape blocks of text containing characters which would otherwise be recognized as markup.

Extensible Markup Language (XML) 1.0 (Fifth Edition)

So I think that a sensible feed-reader should see the CDATA block, grab the HTML inside it, and display it as-is. No need to unescape anything.

That said, I'll see if I can change my feed to not need this hybrid format. There's a brilliant blog post by Suren Enfiajyan which makes the case that regular escaping is probably good enough.

If you've experienced this bug - or think that I'm generating my feeds in the wrong way - please get in touch.