Β· 11 min read

How to Build a Better Time Tracking System With AI


Most time tracking systems still rely on people remembering to start and stop timers, which becomes unreliable as soon as the day gets fragmented across meetings, Slack, code, QA, analytics, and client work.

For developers and agency teams, that fragmentation is normal. Work rarely happens in clean, uninterrupted blocks, and asking someone to manually document every context switch creates more overhead than it should.

A better approach is to treat time tracking as a data collection and classification problem. Your computer already produces most of the signals needed to understand what you are working on. The system can capture those signals automatically, classify the obvious activity with simple rules, and use AI to help with the parts that are harder to interpret.

The result is not fully automated time tracking. It is a timesheet that is mostly assembled before you review it.

Treat time tracking as a classification problem

Traditional time tracking makes the user responsible for nearly everything. You start the timer, choose the project, stop the timer, and write the description.

A more useful system reverses that process. It collects evidence about your activity and tries to determine which project the work belongs to, how long the activity lasted, and how it should be described on a timesheet.

That model fits well with how developers already work because useful project signals are being generated throughout the day whether you track them or not.

Start with automatic activity capture

The first part of the system is a lightweight desktop activity collector.

It can record metadata such as:

  • Active application
  • Window title
  • Browser domain
  • VS Code workspace
  • Git repository
  • Idle state
  • Timestamp

There is no need for AI at this stage. The collector only needs to create a reliable timeline.

A simple timeline could look like this:

TimeActivity
9:00–9:45VS Code β€” client-a-site
9:45–10:15Zoom β€” Client A Weekly
10:15–11:25Chrome β€” Adobe Experience Manager
11:25–11:50Slack
1:00–2:20VS Code β€” client-b-drupal

Even this basic level of capture changes the workflow significantly. Instead of trying to remember what happened at the end of the day, you have something concrete to review.

Classify obvious activity with rules

Once activity is being collected, the next step is to map known signals to projects.

This should be deterministic whenever possible.

For example:

Repository: client-a-site
β†’ Client A
Domain: client-a.com
β†’ Client A
Domain: dashboard.pantheon.io
Site: client-b
β†’ Client B
VS Code workspace: agency-website
β†’ Internal

The same approach can work with Figma projects, Jira boards, CMS environments, hosting dashboards, local development URLs, Slack channels, and application titles.

If a repository or domain clearly belongs to one project, there is little value in asking an LLM to infer it. Simple rules are cheaper, easier to debug, and usually more accurate.

Use calendar data for meetings

Calendar data is another useful source because meetings already come with clean time boundaries and useful context.

A calendar event often includes a title, start and end time, attendees, description, and some indication of the client or project involved.

For example:

9:00–9:30
Client A weekly meeting

followed by:

9:30–11:10
VS Code
Repository: client-a-site

The calendar event gives useful context to the development block that follows it. A Zoom session by itself does not tell you much, but a calendar event called Client A Analytics Review is easy to classify.

This becomes especially useful on days with a lot of short meetings, reviews, and project switching.

Use Git as project evidence

Git activity provides another strong signal for development work.

Commits can give you:

  • Repository
  • Branch
  • Timestamp
  • Commit message
  • Pull request
  • Issue reference

If the system sees VS Code open on client-a-site for an hour and there are commits in the same repository during that period, it has strong evidence that the activity belongs to Client A.

For example:

10:02–11:14
Application: VS Code
Workspace: client-a-site

10:48
Commit: Fix CTA tracking event

11:02
Commit: Remove duplicate analytics trigger

Git should not be used to calculate duration by itself, but it works well as supporting context for both project classification and later description generation.

Group raw activity into work blocks

Individual activity events are too noisy to work with directly.

A raw activity stream could look like this:

10:01 Chrome
10:02 VS Code
10:03 Chrome
10:04 VS Code
10:07 Terminal
10:08 Chrome

Those events are more useful when they are grouped into a block:

10:01–11:12

Applications:
- VS Code
- Terminal
- Chrome

Workspace:
client-a-site

Domains:
- client-a.com
- experience.adobe.com

Git:
- Fix CTA tracking event

The system can create new blocks when the active project changes, a meeting begins, the repository changes, the user goes idle, or there is a long enough gap in activity.

This preprocessing step makes the data much easier to classify and keeps the AI layer focused on useful context rather than raw event noise.

Use AI for ambiguous classification

Once the system has grouped activity and applied known rules, AI can help with the blocks that are still unclear.

A structured activity block might look like this:

Time: 10:00–11:20

Applications:
- VS Code
- Chrome
- Terminal

Workspace:
client-a-site

Domains:
- experience.adobe.com
- client-a.com

Git commits:
- Fix CTA event listener
- Remove duplicate analytics trigger

Previous block:
Client A analytics meeting

The model can then return structured output such as:

{
  "project": "Client A",
  "category": "Development",
  "confidence": 0.97,
  "description": "Implemented and tested analytics tracking and fixed duplicate event firing."
}

This is a much stronger use of AI than asking it to reconstruct an entire day from a vague prompt. The model is working from actual activity data and project context.

Give the model project context

Classification improves when the model knows a small amount about each project.

A project profile could include:

Project: Client A

Repositories:
- client-a-site
- client-a-components

Domains:
- client-a.com
- experience.adobe.com

Tools:
- Adobe Experience Manager
- Adobe Launch
- GA4

Common work:
- Frontend development
- Analytics
- QA
- CMS support

This does not need to be a large knowledge base. A few structured fields are usually enough to help distinguish one client from another.

You can also include project codes, known contacts, Slack channels, Jira projects, and common task types where they help with classification.

Add confidence scores

The system should also make uncertainty visible instead of forcing every block into a project.

A high-confidence block could look like this:

Workspace: client-a-site
Repository: client-a-site
Domain: client-a.com

That can safely be assigned to Client A.

A medium-confidence block might contain Slack and browser research between two Client A blocks. The system can make a reasonable guess while still marking it for review.

A low-confidence block might only contain email, Slack, and general browser activity with no project-specific signals. In that case, leaving it unclassified is better than making up an answer.

Confidence scoring gives you a simple way to decide which blocks can be accepted automatically and which ones should be shown to the user.

Use AI to generate timesheet descriptions

Once the project and duration are known, description generation is one of the more useful places to apply AI.

Raw activity might look like this:

Repository:
client-a-site

Commits:
- Fix event listener
- Update Adobe Launch rule
- Remove duplicate trigger

Domains:
- experience.adobe.com
- staging.client-a.com
- omnibug.io

The system can turn that into:

Implemented and tested analytics tracking, updated Adobe Launch configuration, and resolved duplicate event firing.

The same underlying activity can also be summarized differently depending on where the description will appear.

For internal tracking:

Analytics implementation and QA.

For client billing:

Implemented and tested analytics tracking, updated event handling, and verified tracking behavior in staging.

That keeps the description useful without asking the user to write it manually.

Keep a human review step

I would still keep a review step before submitting anything.

Activity data is useful, but it is not perfect. A client site can stay open while you answer an internal Slack message. VS Code can remain active while you are on a phone call. A browser tab can sit in the foreground while you are away from the computer.

A daily review gives the user a chance to fix those edge cases.

For example:

Client A
3h 15m
Analytics implementation and QA

Client B
2h 40m
Drupal development and content QA

Internal
1h 05m
Planning and project communication

Unclassified
30m
Slack and browser activity

The user only needs to review the parts that look wrong instead of entering the full day from scratch.

Slack can work as the review interface

For teams already working in Slack, a separate dashboard may not be necessary.

A bot could send a draft timesheet near the end of the day:

Today's draft timesheet

Client A β€” 3h 15m
Client B β€” 2h 40m
Internal β€” 1h 05m
Unclassified β€” 30m

The user could reply with corrections such as:

Move 20 minutes from Client A to Internal.

The unclassified block belongs to Client B.

Merge the two Client A entries and describe them as analytics implementation and testing.

Once the corrections are made, the user can approve the final result and send it to the existing time tracking system.

This keeps the review process inside a tool the team is already using.

A practical architecture

A simple implementation could look like this:

Desktop activity collector
          ↓
Local activity store
          ↓
Activity block generator
          ↓
Rule-based classifier
          ↓
Calendar + Git + browser context
          ↓
AI classification
          ↓
AI description generation
          ↓
Human review
          ↓
Harvest / Toggl / Clockify / other API

The order is important because AI works better when it is given structured context.

The collector establishes what happened. Rules handle known mappings. Calendar and Git add supporting evidence. AI deals with ambiguous classification and language.

Collect less data than you think you need

A system like this can become invasive if it collects too much.

In most cases, there is no reason to store screenshots, keystrokes, full Slack messages, email contents, or complete browser history.

Metadata is often enough:

Application: VS Code
Workspace: client-a-site
10:03–10:42

or:

Application: Chrome
Domain: experience.adobe.com
10:42–11:08

That is usually enough to classify the work without capturing the actual content.

Keeping the data limited also makes the system easier to trust if it is eventually used across a team.

Start with a small first version

The first version does not need every integration.

I would start with:

  1. Capture active application, browser domain, workspace, repository, and idle time.
  2. Store events locally.
  3. Map known repositories and domains to projects.
  4. Import calendar events.
  5. Group activity into blocks.
  6. Send ambiguous blocks to an LLM for classification.
  7. Generate a daily summary.
  8. Review the result before submitting the final timesheet.

That is enough to test whether the workflow actually saves time.

Slack integration, Jira or Linear context, GitHub enrichment, and direct timesheet APIs can be added later if the basic model works.

A better use of AI for time tracking

The useful role for AI is not measuring time. Software can already do that reliably.

AI is more useful for interpreting context, resolving ambiguous activity, and turning raw events into descriptions that are understandable to a person reviewing or billing the work.

A practical system should aim for mostly automatic capture with a small amount of human correction. If the system can assemble most of the day accurately and leave only a few uncertain blocks to review, it has already removed most of the annoying part of time tracking.

The computer is already generating most of the information needed to build a timesheet. The implementation challenge is organizing those signals well enough that the final review takes a few minutes instead of requiring someone to reconstruct the entire day.

More Blog Posts

Stay in the loop

You're subscribed!

Check your inbox to confirm your subscription.

About the Author

Want to learn more about my background, experience, and what drives my work? Check out my full story and connect with me.

Learn more about me