cfcollection created an empty container. cfindex is what actually puts data into it, either files on disk, or (the far more common case in a real application) the results of a CFML database query, so they become full-text searchable.
Learning Objectives
After completing this lesson, you'll be able to:
- Index a CFML query's results into a collection with type="custom".
- Index files in a directory with type="path".
- Know the difference between the four index actions: update, delete, purge, and refresh.
- Keep a collection's contents in sync as underlying data changes.
Index Actions
| action | What It Does |
|---|---|
| update | Adds new documents, updates existing ones (matched by key), default |
| delete | Removes specific documents from the index, matched by key |
| purge | Removes everything from the collection, taking it offline momentarily |
| refresh | Purges, then re-indexes, effectively a full rebuild |
A Real Example: Indexing Query Results
This is the workhorse pattern: run a query against your database, then hand its results straight to cfindex so every row becomes a searchable document.
articles = queryExecute("SELECT articleId, title, body FROM articles");
cfindex(
collection = "productDocs",
action = "update",
type = "custom",
query = articles,
key = "articleId",
title = "title",
body = "body"
);key, title, and body here are column names from the query, not literal values, cfindex reads each row and indexes those columns' content for every article.
Indexing Files on a Directory
<cfindex collection="productDocs"
action="update"
type="path"
key="C:\docs\manuals"
recurse="yes"
extensions=".htm,.html,.pdf">type="path" indexes every matching file in that directory (and subdirectories, if recurse="yes"). type="file" indexes one specific file instead, with key pointing at its exact path.
Key Attributes
| Attribute | Meaning |
|---|---|
| collection | Which collection to index into (required) |
| action | update (default), delete, purge, or refresh |
| type | custom (a query's results), file, or path |
| query | The CFML query object, required when type="custom" |
| key | Unique identifier column (custom) or file/directory path (file/path) |
| title | Title column (custom) or a literal title |
| body | Content column(s) to actually index and make searchable |
| recurse | yes/no, include subdirectories (type="path" only) |
| extensions | Comma-separated file extensions to include (type="path") |
Keeping an Index in Sync
A search index isn't automatically updated just because the underlying database changes, it stays exactly as it was until cfindex runs again. A common real pattern: call cfindex action="update" right after any insert or update to the source data, so the collection never drifts far out of sync.
// After inserting/updating an article's row:
cfindex(
collection = "productDocs",
action = "update",
type = "custom",
query = queryExecute("SELECT articleId, title, body FROM articles WHERE articleId = :id", { id: newArticleId }),
key = "articleId",
title = "title",
body = "body"
);Common Beginner Mistakes
Assuming the search index updates itself when the database changes
It doesn't. cfindex has to be called again, with the changed rows, or the index keeps returning stale results indefinitely.
Forgetting recurse="yes" when indexing a directory tree
Without it, type="path" only indexes files directly in that folder, not in any subfolders, easy to miss until search results are mysteriously incomplete.
Passing literal values instead of column names to key/title/body with type="custom"
Those three attributes expect column names from the query object, not the values themselves, cfindex reads them per-row automatically.
Best Practices
- Index query results (type="custom") right after the data that feeds them changes, to keep search results current.
- Use action="refresh" for a full rebuild when a collection's data has drifted significantly, rather than deleting and recreating the collection itself.
- Set extensions explicitly when indexing a directory, to avoid accidentally indexing files that were never meant to be searchable.
Interview Questions
How would you make a database table's rows full-text searchable?
Query the table, then pass that query object to cfindex with type="custom", mapping key, title, and body to the relevant column names. Re-run cfindex whenever the underlying rows change to keep the index current.
What's the difference between cfindex's update and refresh actions?
update adds new documents and updates existing ones (matched by key), leaving everything else in the collection untouched. refresh purges the entire collection first, then re-indexes from scratch, effectively a full rebuild.
Summary
In this lesson, you indexed a CFML query's results into a collection with type="custom", indexed files on disk with type="path", and covered the four index actions (update, delete, purge, refresh) that manage what a collection actually contains over time.
What's Next?
The next lesson covers cfsearch: actually querying the collection you've now created and populated, with relevance-ranked results and highlighted matches.