Personalized Search: Methods for Understanding User Intent 
Personalizing search is hard – because understanding user intent is challenging. Here are techniques to decoding user intent.
Trey Grainger
By Trey Grainger

There are a seemingly endless number of statistics that convey end users’ frustration with search.

  • 71% of websites require users to search by the exact same wording (and/or spelling) the site uses, meaning that they fail to return relevant search results. For example, if the customer searches for “whipped cream” but the website only uses “dessert topping” – null results may appear.
  • Only 1 in 10 people find what they are looking for when using search 
  • More than two-thirds of shoppers say they won’t return to a site with poor search experience

Clearly, we have a ways to go in creating a better search. The key is to understanding user intent. There are three types of contexts that come into play to understand user intent: 

  1. content 
  2. domain 
  3. user 

Understanding the user context is a particularly hard problem. We like to believe that audiences represent a handful of personas. Indeed, signals-boosting models find the most popular answers across all users.

Unfortunately, personas don’t search, people do. Because of language, experience, and or cultural issues, signals boosting alone might not be the best answer. Personalized search instead attempts to learn about each specific user’s interests and to return search results catering to those interests.

The key to personalized search is to leverage user signals to learn latent features describing users’ interests. You can use these signals/ latent features and lexalytics (keyword search) to personalize search and return search results catering to those interests.

But first, let’s take a step back and remember that keyword search represents only content understanding, and recommendation engines typically accept no direct user input and recommend on inferred knowledge. This is called collaborative filtering – more on that in a bit. 

Both content understanding and user understanding should be combined when possible. 

Personalized search lies at the intersection between keyword search and collaborative recommendations. There is a broad spectrum of capabilities that lie within the personalization spectrum between search and recommendation systems. 

What Are Latent Features?

In general, latent features refer to hidden characteristics. These might be characteristics that are not directly observable, but can be inferred by data patterns, relationships, or statistical methods. 

For example, when searching for restaurants, a user’s location clearly matters. When searching for a job, prior employment history (previous job titles, experience level, salary range) and location may matter. When searching for products, particular brand affinities, colors of appliances, complementary items purchased, and similar personal tastes may matter.

These latent features can be used to generate product recommendations and boosts to personalized search results. 

We can also use content-based embeddings to relate products and can leverage embeddings of the products each user interacts with to generate vector-based personalization profiles to personalize search results. 

An embedding is a numerical vector (usually a list of floats) that is intended to represent the semantic meaning of a given term sequence. Embeddings can represent term sequences of any length, but when representing individual words or phrases, we call the embeddings word embeddings. If representing products it would be product embeddings, etc.) 

You index embeddings, so that when a person searches, that query is also turned into a vector. Answers are ranked by how semantically similar the query is to all your embeddings. In practice,  if a person searches on, and selects tennis shoes, and then is looking for socks, we might order by athletic socks.

You can cluster products by their embeddings to generate personalization guardrails to ensure that users don’t see personalized search results based on products from unrelated categories.

Leveraging Latent Features in Personalized Search

When searching for restaurants, a user’s location clearly matters. When searching for a job, each user’s employment history (previous job titles, experience level, salary range) and location may matter. When searching for products, particular brand affinities, colors of appliances, complementary items purchased, and similar personal tastes may matter. Let’s look at how latent features may impact search.

Personalized queries

Let’s imagine we’re running a restaurant search engine. Our user, Michelle, is on her phone in New York at lunchtime, and she types in a keyword search for steamed bagels. She sees top-rated steamed bagel shops in Greenville, South Carolina (USA), Columbus, Ohio (USA), and London (UK).

What’s wrong with these search results? Well, in this case, the answer is clear — Michelle is looking for lunch in New York, but the search engine is showing her results hundreds to thousands of kilometers away. 

Michelle never told the search engine she only wanted to see results in New York, nor did she tell the search engine that she was looking for a lunch place close by because she wants to eat now. Nevertheless, the search engine should be able to infer this information and personalize the search results accordingly.

Consider another scenario — Michelle is at the airport after a long flight, and she searches on her phone for driver. The top results that come back are for a golf club for hitting the ball off a tee, followed by a link to printer drivers, followed by a screwdriver. If the search engine knows Michelle’s location, shouldn’t it be able to infer her intended meaning — that she is searching for a ride?

Using our job search example from earlier, let’s assume Michelle goes to her favorite job search engine and types in nursing jobs. Like our restaurant example earlier, wouldn’t it be ideal if nursing jobs in New York showed up at the top of the list? What if she later types jobs in Seattle? Wouldn’t it be ideal if — instead of seeing random jobs in Seattle (doctor, engineer, chef, etc.) — nursing jobs now showed up at the top of the list, since the engine previously learned that she is a nurse?

Each of these is an example of a personalized query: the combining of both an explicit user query and an implicit understanding of the user’s intent and preferences into a search that serves results specifically catering to that user. Doing this kind of personalized search well is tricky, as you must carefully balance your understanding of the user without overriding anything they explicitly want to query. When it’s done well, though, personalized queries can significantly improve search relevance.

User-guided recommendations

Just as it’s possible to sprinkle an implicit understanding of user-specific attributes into an explicit keyword search to generate personalized search results, it’s also possible to enable user-guided recommendations by allowing user-overrides of the inputs into automatically generated recommendations.

It is becoming increasingly common for recommendation engines to allow users to see and edit their recommendation preferences. These preferences usually include a list of items the user interacted with before by viewing, clicking, or purchasing them. 

Across a wide array of use cases, these preferences could include both specific item preferences, like favorite movies, restaurants, or places, as well as aggregated or inferred preferences, like clothing sizes, brand affinities, favorite colors, preferred local stores, desired job titles and skills, preferred salary ranges, and so on. 

These preferences make up a user profile: they define what is known about a customer, and the more control you can give a user to see, adjust, and improve this profile, the better you’ll be able to understand your users and the happier they’ll likely be with the results.

Recommendation Algorithm Approaches

In general, recommendation engine implementations are based on what data is available to drive their recommendations. Some systems only have user behavioral signals and very little content or information about the items being recommended, whereas other systems have rich content about items, but very few user interactions with the items. 

Here’s a review of content-based, behavior-based, and multimodal recommenders.

Content-based recommenders

These algorithms recommend new content based on attributes shared between different entities. This can be between users and items, between items and items, or between users and users. 

For example, imagine a job search website. Jobs may have properties on them like “job title”, “industry”, “salary range”, “years of experience”, and “skills”. Users will have similar attributes on their profile or resume/CV. 

Based upon these properties, a content-based recommendation algorithm can figure out which of these features are most important and can then rank the best matching jobs for any given user based on the user’s desired attributes. This is what’s known as a user-item (or user-to-item) recommender.

Similarly, if a user likes a particular job, it is possible to leverage this same process to recommend similar jobs based on how well those jobs match the attributes of the first job. This type of recommendation is popular on product details pages, where a user is already looking at an item and it may be desirable to help them explore related items. This kind of recommendation algorithm is known as an item-item (or item-to-item) recommender.

Figure 9.3 demonstrates how a content-based recommender might leverage attributes about items with which a user has previously interacted to match similar items for that user. In this case, our user viewed the “detergent” product and was then recommended “fabric softener” and “dryer sheets” based upon these items matching within the same category field (the “laundry” category) and containing similar text to the “detergent” product within their product descriptions.

NOTE: It’s also possible to match users to other users, or any entity to any other entity. In the context of content-based recommenders, all recommendations can be seen as item-item recommendations, where each item is an arbitrary entity that shares attributes with the other entities being recommended.

Behavior-based recommenders

Behavior-based recommenders leverage user interactions with items (documents) to discover similar patterns of interest among groups of items. This process is called collaborative filtering, referring to the use of a multiple-person (collaborative) voting process to filter matches to those demonstrating the highest similarity, as measured by how many overlapping users interacted with the same items. 

The idea here is that similar users (i.e., those with similar preferences) tend to interact with the same items, and when users interact with multiple items, they are more likely to be interacting with similar items as opposed to unrelated items.

Collaborative filtering algorithms fully crowdsource the relevance scoring process from your end users. In fact, features of the items themselves (name, brand, color, text, and so on) are not needed — all that is required is a unique ID for each item and knowledge of which users interacted with which items. 

Further, the more user-interaction signals you have, the smarter these algorithms tend to get, because more people are continually voting and informing your scoring algorithm. This often leads to collaborative filtering algorithms significantly outperforming content-based algorithms.

The image above demonstrates how overlapping behavioral signals from multiple users can be used to drive collaborative recommendations. In this figure, a new user is expressing interest in fertilizer, and because other users who have previously expressed interest in fertilizer tend to also click on, add to cart, or purchase soil and mulch, then soil or mulch will be returned as recommendations. 

Another behavior-based cluster of items including a screwdriver, hammer, and nails is also depicted, but they don’t sufficiently overlap with the user’s current interest (fertilizer), so they are not returned as recommendations.

NOTE: Unfortunately, the same dependence upon user behavioral signals that makes collaborative filtering so powerful also turns out to be its weakness. What happens when there are only a few interactions with a particular item — or possibly none at all)?

The answer is that the item either never gets recommended (when there are no signals), or it will be likely to generate poor recommendations or show up as a bad match for other items (when there are few signals). This situation is known as the cold-start problem, and it’s a major challenge for behavior-based recommenders. To solve this problem, you typically need to combine behavior-based recommenders with content-based recommenders.

Multimodal recommenders

Multimodal recommenders (also known as hybrid recommenders) combine both content-based and behavior-based recommender approaches. Since collaborative filtering tends to work best for items with many signals, but works poorly when few or no signals are present, use content-based features as a baseline and then layer a collaborative filtering model on top. 

This way, if few signals are present, the content-based matcher will still return results, whereas if there are many signals, the collaborative filtering algorithm will take greater prominence when ranking results. 

Image of how content filtering and collaborative filtering work to understand user intent.

You can see that the user could interact with either the drill (which has no signals) or the screwdriver (which has previous signals from other users, as well as content), and the user would receive recommendations in both cases. This provides the benefit that signals-based collaborative filtering can be used, while also enabling content-based matching for items with insufficient signals.

Multimodal recommendations combine both content-based matching and collaborative filtering into a hybrid matching algorithm. Incorporating can give you the best of both worlds: high-quality crowdsourced matching, while avoiding the cold-start problem for newer and less-well-discovered content. 

Out of all the techniques in AI-powered search, personalization is both one of the most underutilized ways to better understand user intent and one of the most challenging. 

While recommendation engines are prevalent, the personalization spectrum between search and recommendations is more nuanced and less explored. So long as personalized search is implemented with care, it can be a powerful tool to drive more relevant search results and save the user time discovering the items that best meet their particular interests.

Similar Posts

Sign Up

SIGN UP AND RECEIVE EXCLUSIVE UPDATES