AI SEARCH TOOL

USABILITY EVALUATION

Analyzing the ROI of AI-powered search for prospective students.

Lead UX Researcher | Georgia Institute of Technology | March 2026

When Georgia Tech's Institute Communications asked me to evaluate an AI search tool on our website ahead of an April 2026 renewal deadline, the stakes were clear: justify the investment or recommend walking away. What I discovered wasn't just about one tool's shortcomings, it revealed fundamental truths about how students interact with AI, what drives trust versus adoption, and why familiarity often wins over functionality.

The business question was straightforward but high-stakes: Could this AI search tool justify its subscription cost by becoming students' primary resource for exploring majors and degrees? Or would it remain a nice-to-have supplement to tools like Google and ChatGPT that prospective students already use for free?

Institute Communications entered into an agreement with the vendor knowing we were adopting an early-stage product, meaning the product had never been pushed live with any other organization. With no benchmark data and no comparable user research to draw from, we were starting from scratch. That context shaped everything about how I approached the evaluation. Rather than running a single usability study, I designed a three-phase research process: first auditing the interface on its own terms, then isolating AI performance from UI friction, and finally measuring real-world adoption and user behavior.

Phase 1

UX Audit

The first phase included a UX Audit in which I performed a heuristic evaluation to identify interface friction and technical bugs before launching user tests. This "clean house" phase helped distinguish surface-level design flaws from deeper AI performance issues. The audit revealed critical discoverability issues, with no placeholder text or sample queries to guide users. I also identified a major gap in system feedback: the interface lacked a loading indicator or any visual confirmation that a search was in progress. Without this, users were left in the dark during the time it took for the AI to generate results.

The original search bar without placeholder text or an "X" button.

I identified several other interaction flaws, such as the absence of an "X" button to reset searches and significant mobile issues, including overlapping icons and a keyboard that failed to dismiss after a query. A critical finding first identified in this audit involved error handling: when the tool failed to find an answer, the "No Results" screen offered no alternative paths or related topics to keep the user engaged. Finally, the information architecture made results difficult to scan, as long AI summaries were presented as a "wall of text" without necessary line breaks.

Phase 2

Intervention with High-Fidelity Prototyping

Following the heuristic audit, I used the Georgia Tech Brand Guidelines to build a high-fidelity Figma prototype. The goal was to isolate whether user frustration stemmed from a lack of UI polish or the core AI performance. While the results section remained a fixed third-party component, my redesign focused on the elements within our control:

System Feedback & User Control: I implemented a dedicated "Clear" (X) button within the search input and a persistent "Back to Degrees" breadcrumb. This addressed a critical friction point where users felt "trapped" in a search state, providing a clear exit strategy and reducing interaction cost.

Input Optimization: I refined the search bar for brand alignment, ensuring the tool felt like a native part of the Georgia Tech ecosystem.





By pushing this high-fidelity version live, I ensured that the subsequent usability testing would measure the AI’s value proposition, rather than surface-level usability bugs.

Phase 3

Quantitative Baseline & Behavioral Analysis

Following the initial audit, I transitioned to a data-driven phase to measure real-world adoption. By analyzing usage logs and web analytics from February to March, I established a performance baseline that revealed a low 9% adoption rate among the target audience and a 26% follow-up search rate, indicating that even users who tried the tool often needed to search again because their initial results were insufficient.

Comparative Usability Testing

While the analytics told us what was happening, I needed to speak with prospective students to understand why. I conducted moderated usability sessions with six prospective students to evaluate trust, ease of use, and long-term viability. I used a comparative framework, giving prospective students tasks—such as finding specific degree requirements—to complete using both Searchify and their preferred tools like Google or ChatGPT. The sessions revealed several critical friction points:

  1. Content & Accuracy: The AI summaries often lacked depth or returned inaccurate results. In one instance, a search for "media" yielded "No Results" on Searchify, while a Google search for the same term immediately surfaced Georgia Tech's Literature, Media, and Communication degree.
  2. UI Scannability: Participants struggled with the visual hierarchy, finding the "wall of text" results difficult to parse and the purpose of the citations unclear.
  3. The "Brand Halo" Effect: Interestingly, users did trust the tool, but not because of its performance. Their trust was tied entirely to the Georgia Tech digital ecosystem. As one participant noted, because the tool was on the official site, they assumed the information was coming directly from the source.

Our primary finding was that Searchify could not function as a standalone tool. Participants were split: half would not switch from their original search methods, while the other half saw the AI search tool only as a supplemental resource rather than a replacement. This insight moved the conversation beyond UI fixes and directly addressed the core business challenge: whether the tool provided enough unique value to justify the subscription cost.

Beyond the Primary Question: Strategic Insights for Future AI Development

How Interface Design Shapes User Behavior

Design-Driven Prompting: I observed a stark difference in how students interacted with different interfaces. While they asked rich, conversational questions in chat-style tools, Searchify's search-box design "trained" them to use simple keywords. This effectively lowered the quality of the AI's output because users weren't providing the context necessary for the tool to succeed.
Natural Interaction Style: Overall, participants prefer conversational style and writing longer, more specific questions when searching for information.

What Drives Verification

AI-Native Verification: AI-native users take summaries and results with a grain of salt and want to verify big decisions with original sources.

The Power of Familiarity

Wait-Time Tolerance: Despite stakeholder concerns regarding response latency, participants were remarkably patient with AI generation. One user noted that "taking a little bit to index" actually made sense for an AI tool. This shifted our focus: speed wasn't the dealbreaker—the quality of the results was.
The "Familiarity" Barrier: The toughest competitor wasn't a specific feature, but the deep-seated habits users had with Google and ChatGPT. One participant admitted they simply knew "what keywords to put in" to get what they wanted from familiar tools, highlighting that any new AI tool must provide massive, obvious value to overcome the cost of switching workflows.

The research concludes that the tool fails to meet the core needs of prospective students, who require a level of content depth that the current iteration cannot provide. While the Georgia Tech brand offers a baseline of trust, users consistently reported that they would continue to rely on familiar, established search methods like Google or ChatGPT. Because the tool would only serve as a supplemental—rather than primary—resource, the subscription cost cannot be justified against current business goals or user adoption metrics.

Future AI search tools must transition to chat-style interfaces to align with conversational mental models while prioritizing source transparency through prominent, accurate, and easily accessible verification links. While AI summaries offer potential value, maintaining the current Google-powered search without Vertex AI integration remains the most viable path due to budget constraints and the accessibility of superior free alternatives.