DevLearningTools

MODULE 15 · LESSON 02

XSS Prevention

Cross-site scripting is the output-side mirror of SQL injection: encoding untrusted data correctly for the exact context it lands in, whether that's HTML, an attribute, JavaScript, CSS, or a URL, plus sanitizing rich text and hardening cookies.

New lessons are added one at a time as the course gets built out — a graded quiz for each lesson is still on the way.

The last lesson was about the input side: never letting untrusted data become part of a SQL query. XSS prevention is the same idea applied to output: never letting untrusted data become part of the page's HTML, JavaScript, or CSS without first encoding it for exactly the spot it's landing in.

Learning Objectives

After completing this lesson, you'll be able to:

  • Explain why the correct encoding function depends on the output context, not just "is this user input."
  • Use encodeForHTML(), encodeForHTMLAttribute(), encodeForJavaScript(), encodeForCSS(), and encodeForURL() correctly.
  • Use the <cfoutput encodefor="..."> shortcut to encode an entire block at once.
  • Sanitize rich text/HTML input with getSafeHTML() and isSafeHTML() instead of breaking it with plain encoding.
  • Harden session cookies with httpOnly and secure as a defense-in-depth layer.

How Context-Aware Encoding Fits Together

Untrusted Data Arrives

form, url, or cookie value

Identify the Output Context

HTML body, attribute, script, CSS, URL

Apply the Matching Encoder

encodeForHTML(), encodeForJavaScript()...

Safely Rendered

no script execution possible

Why Context Matters

The same raw string is dangerous in different ways depending on where it ends up. A value placed into an HTML body only needs its angle brackets escaped. The same value placed inside an HTML attribute, inside a <script> block, inside a CSS rule, or inside a URL query string each has a different set of characters that can break out of that context, so each context needs its own encoding function.

Output ContextFunctionExample
HTML bodyencodeForHTML()<div>#encodeForHTML(url.username)#</div>
HTML attributeencodeForHTMLAttribute()<input value="#encodeForHTMLAttribute(form.val)#">
Inside <script>encodeForJavaScript()<script>var name = "#encodeForJavaScript(form.name)#";</script>
CSS ruleencodeForCSS()background-color: #encodeForCSS(url.color)#;
URL query stringencodeForURL()search.cfm?q=#encodeForURL(form.query)#
NOTE

All five take the same two arguments: value (required) and canonicalize (optional boolean, default false, normalizes mixed/double-encoded input before encoding).

The Shortcut: cfoutput's encodefor Attribute

Instead of wrapping each variable individually, encodefor on <cfoutput> (CF2016+, Lucee 5.1+) applies one encoder to every variable rendered inside that block.

Tag Syntax
<cfoutput encodefor="html">
    <p>Welcome back, #session.username#</p>
</cfoutput>
NOTE

Adobe's accepted values are html, htmlattribute, javascript, css, xml, xmlattribute, url, xpath, ldap, and dn. Lucee documents the same idea with slightly different spelling for the multi-word contexts (html_attr, xml_attr) and adds vbscript, so an encodefor value copied between the two engines is worth double-checking against that engine's own docs.

Rich Text Input: Sanitize, Don't Encode

A comment box or a WYSIWYG editor is a case where the user is meant to submit real HTML. Encoding it would escape every tag and break the formatting. The fix here is sanitizing it instead, stripping out anything dangerous (script tags, inline event handlers like onerror or onload) while keeping the safe markup intact.

FunctionReturnsPurpose
getSafeHTML(inputString [, policyFile, throwOnError])StringStrips unsafe markup out of inputString using AntiSamy policy rules, returning the cleaned HTML
isSafeHTML(inputString [, policyFile])BooleanChecks inputString against the same policy rules without modifying it
Tag Syntax
<cfset cleanComment = getSafeHTML(form.userComment) />
<cfoutput>#cleanComment#</cfoutput>

A Real Detail: This Pair Isn't Identical Across Engines

getSafeHTML() and isSafeHTML() are built on Adobe ColdFusion's bundled OWASP AntiSamy integration, and they're documented on cfdocs.org as straightforward ColdFusion functions. Lucee's own documentation site, by contrast, has no equivalent page for either function. Even encodeForHTML() itself, on Lucee's own docs, is listed as requiring Lucee's separate Guard extension to be installed, it isn't assumed to just be there the way it is on Adobe ColdFusion. If a project needs to run on both engines, verify these specific functions exist on the target server rather than assuming parity.

Defense in Depth: Cookie Hardening

Encoding output correctly is the main defense, but if an XSS payload ever did slip through, a hardened session cookie limits what it can do. cfcookie's httpOnly attribute (CF9+) stops JavaScript from reading the cookie via document.cookie at all, and its secure attribute stops the cookie from ever being sent over a plain, unencrypted HTTP connection.

Tag Syntax
<cfcookie name="CFID" value="#session.sessionid#" httponly="true" secure="true">

Common Beginner Mistakes

Using encodeForHTML() everywhere, regardless of context

encodeForHTML() only escapes characters that are dangerous in an HTML body. Used inside a <script> block or a CSS rule, it leaves that context's actual dangerous characters untouched. Match the function to the context, not just to "this is user input."

Encoding rich text input instead of sanitizing it

encodeForHTML() on a WYSIWYG editor's output escapes every tag the user intentionally added, breaking the formatting entirely. That case calls for getSafeHTML(), which strips only the dangerous parts.

Assuming cookie hardening replaces output encoding

httpOnly and secure reduce what a successful XSS payload can do, they don't stop the injection from happening in the first place. They're a backstop, not a substitute for encoding output correctly.

Best Practices

  • Pick the encodeFor* function that matches the exact output context, not a single default used everywhere.
  • Use getSafeHTML()/isSafeHTML() for intentional rich-text input, never plain encoding.
  • Set httpOnly and secure on session cookies as a defense-in-depth layer, not as the primary fix.
  • Verify encodeForHTML(), getSafeHTML(), and encodefor's exact accepted values against the target engine's own docs if the application needs to run on both Adobe ColdFusion and Lucee.

Interview Questions

Why isn't one encoding function enough for every output context?

Each context (HTML body, HTML attribute, inside <script>, CSS, URL) has its own set of characters that can break out of that context. A function built for one context, like encodeForHTML(), doesn't escape the characters that are dangerous in a different context, like inside a <script> block.

When would you use getSafeHTML() instead of encodeForHTML()?

When the input is meant to contain real HTML, like a rich text editor's output. encodeForHTML() would escape every tag and break the formatting, getSafeHTML() strips only the dangerous parts (script tags, inline event handlers) while keeping the safe markup intact.

What does setting httpOnly on a session cookie actually protect against?

It stops JavaScript from reading that cookie via document.cookie. If an XSS payload does get injected, httpOnly prevents it from exfiltrating the session token through that cookie.

Summary

In this lesson, you matched each output context (HTML body, HTML attribute, JavaScript, CSS, URL) to its correct encodeFor* function, used the <cfoutput encodefor="..."> shortcut, sanitized rich text input with getSafeHTML()/isSafeHTML() instead of encoding it, saw a real cross-engine difference in how Adobe ColdFusion and Lucee support these functions, and hardened session cookies with httpOnly and secure as a defense-in-depth layer.

What's Next?

The next lesson covers password hashing: never storing a password as plain text, and why a general-purpose hash function like MD5 or SHA isn't the right tool for the job either.