DevLearningTools

MODULE 14 · LESSON 09

cfindex

Populating a search collection with cfindex: indexing a CFML query's results (the most common real use case), indexing files on disk, and the four actions (update, delete, purge, refresh) that manage what's in the index.

New lessons are added one at a time as the course gets built out — a graded quiz for each lesson is still on the way.

cfcollection created an empty container. cfindex is what actually puts data into it, either files on disk, or (the far more common case in a real application) the results of a CFML database query, so they become full-text searchable.

Learning Objectives

After completing this lesson, you'll be able to:

  • Index a CFML query's results into a collection with type="custom".
  • Index files in a directory with type="path".
  • Know the difference between the four index actions: update, delete, purge, and refresh.
  • Keep a collection's contents in sync as underlying data changes.

Index Actions

actionWhat It Does
updateAdds new documents, updates existing ones (matched by key), default
deleteRemoves specific documents from the index, matched by key
purgeRemoves everything from the collection, taking it offline momentarily
refreshPurges, then re-indexes, effectively a full rebuild

A Real Example: Indexing Query Results

This is the workhorse pattern: run a query against your database, then hand its results straight to cfindex so every row becomes a searchable document.

CFScript
articles = queryExecute("SELECT articleId, title, body FROM articles");

cfindex(
    collection = "productDocs",
    action = "update",
    type = "custom",
    query = articles,
    key = "articleId",
    title = "title",
    body = "body"
);
NOTE

key, title, and body here are column names from the query, not literal values, cfindex reads each row and indexes those columns' content for every article.

Indexing Files on a Directory

Tag Syntax
<cfindex collection="productDocs"
    action="update"
    type="path"
    key="C:\docs\manuals"
    recurse="yes"
    extensions=".htm,.html,.pdf">
NOTE

type="path" indexes every matching file in that directory (and subdirectories, if recurse="yes"). type="file" indexes one specific file instead, with key pointing at its exact path.

Key Attributes

AttributeMeaning
collectionWhich collection to index into (required)
actionupdate (default), delete, purge, or refresh
typecustom (a query's results), file, or path
queryThe CFML query object, required when type="custom"
keyUnique identifier column (custom) or file/directory path (file/path)
titleTitle column (custom) or a literal title
bodyContent column(s) to actually index and make searchable
recurseyes/no, include subdirectories (type="path" only)
extensionsComma-separated file extensions to include (type="path")

Keeping an Index in Sync

A search index isn't automatically updated just because the underlying database changes, it stays exactly as it was until cfindex runs again. A common real pattern: call cfindex action="update" right after any insert or update to the source data, so the collection never drifts far out of sync.

CFScript
// After inserting/updating an article's row:
cfindex(
    collection = "productDocs",
    action = "update",
    type = "custom",
    query = queryExecute("SELECT articleId, title, body FROM articles WHERE articleId = :id", { id: newArticleId }),
    key = "articleId",
    title = "title",
    body = "body"
);

Common Beginner Mistakes

Assuming the search index updates itself when the database changes

It doesn't. cfindex has to be called again, with the changed rows, or the index keeps returning stale results indefinitely.

Forgetting recurse="yes" when indexing a directory tree

Without it, type="path" only indexes files directly in that folder, not in any subfolders, easy to miss until search results are mysteriously incomplete.

Passing literal values instead of column names to key/title/body with type="custom"

Those three attributes expect column names from the query object, not the values themselves, cfindex reads them per-row automatically.

Best Practices

  • Index query results (type="custom") right after the data that feeds them changes, to keep search results current.
  • Use action="refresh" for a full rebuild when a collection's data has drifted significantly, rather than deleting and recreating the collection itself.
  • Set extensions explicitly when indexing a directory, to avoid accidentally indexing files that were never meant to be searchable.

Interview Questions

How would you make a database table's rows full-text searchable?

Query the table, then pass that query object to cfindex with type="custom", mapping key, title, and body to the relevant column names. Re-run cfindex whenever the underlying rows change to keep the index current.

What's the difference between cfindex's update and refresh actions?

update adds new documents and updates existing ones (matched by key), leaving everything else in the collection untouched. refresh purges the entire collection first, then re-indexes from scratch, effectively a full rebuild.

Summary

In this lesson, you indexed a CFML query's results into a collection with type="custom", indexed files on disk with type="path", and covered the four index actions (update, delete, purge, refresh) that manage what a collection actually contains over time.

What's Next?

The next lesson covers cfsearch: actually querying the collection you've now created and populated, with relevance-ranked results and highlighted matches.