Taxwell / AI & People operations

HR inbox agent

After staffing reductions, more of our general HR inbox was landing with me. Some questions were routine, some belonged somewhere else, and some genuinely needed a person to make a judgment. I built a ChatGPT Enterprise agent to help sort out which was which.

It used approved information to answer routine questions, route requests, and flag cases that needed an HR colleague.

My part

I built and maintained the agent, defined the handling rules and information boundaries, and evaluated its behavior using example requests.

The outcome

Deployed; estimated 5–10 team hours per week returned to other HR work. Requests needing judgment stayed with people.

Try the handling examplesSee when the agent answers, routes, or asks for human review.
Example requests
Example requests

Incoming request

Where is the address-change form?

Available information / constraints

A fictional directory lists the form under Forms in the employee portal.

Decision: Answer

Find the form under Forms in the employee portal.

Reason: The available reference answers the question.

Incoming request

I can’t sign into the employee portal.

Available information / constraints

A fictional directory assigns access problems to IT support. The agent cannot diagnose the account.

Decision: Route

Direct the request to IT support.

Reason: The directory identifies the responsible team.

Incoming request

Does the usual rule apply to my exception?

Available information / constraints

A fictional reference describes the general rule but does not cover this exception.

Decision: Human review

Ask an HR colleague to review.

Reason: The information is insufficient for this judgment.

Fictional requests and references. No employee data or live agent.

More about the buildImplementation details and decisions.

Connections and memory

I worked with GRC and system owners on what the agent could access and how it could handle employee information. Some connections weren’t available, so I built around the approved sources and tools, including Outlook and SharePoint.

The built-in memory wasn’t enough for the range of work. Outlook was one of the systems I was allowed to use, and I realized drafts could give the agent a persistent place to store and retrieve context. I used separate drafts for different kinds of memory rather than relying on whatever happened to remain in the active chat.

Testing and maintenance

I reviewed responses to past and test inquiries, scored examples, and changed the instructions, skills, and memory based on what I found. I checked the answer itself, whether the request had gone to the right place, and whether a person should have been involved.

The team estimated that the agent returned about 5–10 hours a week to other HR work. This was a deployed workflow that I refined through examples and review; I wasn’t fine-tuning the model’s weights.