What it took to study 40 million conversations with a generalist LLM

How a researcher turned rare access to Copilot’s usage data in 2025 into a usage story and eventually a paper published in Nature Health.

In December last year, we did something unprecedented not just for the company but in the field: we released a usage report on a generalist LLM using a dataset of over 40 million conversations (It’s About Time: The Copilot Usage Report 2025).

This was, at the time, 10 times bigger than the largest study to date.

The data gave us the first glimpse into how people use Copilot, Microsoft’s consumer chatbot, in their routines, and how these routines propagated through millions of datapoints. The report spread through professional circles and even generalist media.

But you may be wondering: how did we end up with this treasure trove?

It’s not every day that an academic stumbles upon a dream dataset – data that truly shows how people use a product that is so widely spread, with hundreds of millions of active users all over the world. But this was not just about access to the dataset but about having the chance and freedom to investigate. So – let’s take a step back and have a look at all the domino pieces that brought us here.

How it started

It’s been a little over a year since I joined Microsoft AI as a researcher. I was recruited to join an experimental and brand new team, at the time called Advanced Planning Unit – now a much more descriptive “Futures Team” – where the idea was to research the broad impact of AI in society, in the long term.

As a new team, one of our main goals was to meet and engage with other working groups across the company, to understand the intricacies and the incredible work that happens behind the scenes that we don’t get to learn about when we’re users of the products.

In my second week on the job, when I flew from London to San Francisco (my first time travelling outside of Europe, at that!), I met the team that was developing a way of evaluating Copilot’s performance in the wild. They were generating this gigantic pool of information every day, in order to make sure we know that Copilot is working its best for the users.

It made sense that we explored ways we could work together – there were a lot of ideas on how we could collaborate, and what we could investigate – because, let’s be honest, the only way we can infer how people will be relating to AI in the future is if we understand how that relationship is happening right now.

Because, let’s be honest, the only way we can infer how people will be relating to AI in the future is if we understand how that relationship is happening right now.

So, we found our first project goals: to establish access to the Copilot dataset, define research questions and begin structural analysis of how people use Copilot in their daily lives.

So much information

First things first: access required strict governance, diverging from the usual academic protocols. Given the scale of the company and the commitment to protecting user privacy, the data pipeline removed personal identifiers and transformed logs into LLM-generated summaries with predefined topic and intent labels before any analysis occurred. Once access was granted, I began weekly exploratory sessions alongside one of the team members. There was just SO MUCH INFORMATION.

Illustration of a worried woman in pink, flailing in water with graphs, charts, and a life preserver floating around her, symbolizing feeling overwhelmed.

But among peers, as we’re all a bit nerdy in this part of the internet, let me confess: it was incredible! Every week, working closely with the team, I was shaping a dashboard, slicing the data by time, models, devices, anything that I thought could provide a good insight. I got to look at all these patterns, and play around, and really, truly try to grasp how people were relating to this tool, with a dataset that I never, in my academic dreams, would’ve been able to play with.

This is when my pattern-recognition self started to see some interesting results. We saw that the human frequencies were translated into this dataset: users programmed more during the week and gamed more on weekends, used their phones for more personal queries than their laptops. We shared these internally, but the science communicator in me had the will to come out and work on compiling these results into a report that we could share externally.

A woman in a red dress walks among large blue and white speech bubbles scattered across a yellow background, symbolizing navigating conversations or messages.
The post exploded – and I found myself having 15 minutes of LinkedIn fame (a very specific type of fame, let’s be honest).

Bear in mind, at this point, that I had never published anything on behalf of a company, let alone such a large one. My previous work as an academic was very small scale, so the only people who ever had eyes on the papers before they came out were the authors involved. This time, however, it was fascinating: by the time the preprint was released, more eyes had been on the paper and blog post than all of my previous work combined.

The extent to which the work is combed through in order to keep the users protected was remarkable. There was due diligence at every step of the way! Also, for the first time ever, I found myself having meetings at all hours of the day to be able to overlap with my colleagues on the West Coast of the US. After a month of back and forth, of fine-tuning words and graphs, the paper went live. The truth of the matter is that this was an absolute first, for us and for the company, and we didn’t know how it would be received.

The post exploded – and I found myself having 15 minutes of LinkedIn fame (a very specific type of fame, let’s be honest).

A pink hourglass filled with blue and yellow thumbs-up icons sits on a yellow surface against a blue background.

But the work and, most importantly, the questions didn’t stop there. Being a scientist means questioning everything until you have enough answers (spoiler alert: we never do).

One of the main things that we saw, that we didn’t know before, was that health was the top topic on mobile, at any time of any day in 2025. This, along with the serendipity of my background being in health, made me want to unravel more of that yarn. So, with the health team, I started looking deeper and deeper into how people were using Copilot specifically for health.

Copilot, specifically for health

This was a whole other whale! First, to really understand the minutiae of these conversations, while maintaining maximum privacy, the health team developed a set of extra health topics and intents by which to categorize conversations.

With this new taxonomy, where we look into exactly how people use Copilot for health (with categories such as Information and Education, Healthcare Navigation, Symptom Questions and others), we then looked at all the health conversations caught by the daily random sampler in January 2026. We wanted to further challenge ourselves, and instead of releasing just a report, we worked on getting it peer reviewed. And alas – we got it. But for an entire month that’s all I was living and breathing (again). These papers, especially closer to the publication dates, took over my life.

While this blog post was in my drafts, going around in my mind, two more achievements were unlocked. First, our health usage paper made it to the cover of the July issue of Nature Health. Never in my life did I think something like this would happen, for which I’m incredibly grateful.

And secondly, a more extensive piece of work was also published. While we were looking at the reports, we wanted to dig deeper into how each country related to these intents and topics – but this is a more sensitive matter. We did an ethics review and re-checked all the data privacy and compliance steps we could follow to make sure we did right by our users.

But we managed to publish our follow-up paper in Nature Health, looking at a global analysis of country-level factors associated with health queries. This work was some of the most fun I’ve had while on the job – we were backing all of our words with solid statistical evidence, and it was fascinating to see how the world is represented in our dataset.

The corporate life lessons that I’ve learned are endless. But it’s a price worth paying when the reward is the research that we got to do. My final question, then, is: how did the academic in me shape the way this research went?

During my biomedical engineering undergrad back in Portugal, I learned to do as much as possible with the little resources that I had. You don’t actually know how valuable time is until that’s the only thing you have for your work. And sometimes, you’d have none at all.

You don’t get the degree by being the best, working the hardest or being the smartest. You get it by not giving up – so I never did.

My time as a PhD student in Manchester taught me about resilience, especially with COVID having hit the later tail of the degree. But the most important thing I was told is that you don’t get the degree by being the best, working the hardest or being the smartest. You get it by not giving up – so I never did.

Being a Post-doc showed me the life beyond the (computational) lab, where communication is key. Your research can be the best in the world, but if you can’t package it to show to anyone that you know, then you really don’t understand it yet. Everyone should have a right and a chance to learn, and it’s on you, when you know what you want, to actually make sure people can access your knowledge.

So, it wasn’t just the access to the data. It was the freedom to look and question, and it was the mindset of the academic who is hungry for knowledge, with the curation of the science communicator and the willpower to believe that science is better when we share with the world.

Smiling woman with shoulder-length wavy hair wearing a hoodie; sketched lines in the background add a playful effect. Image is in sepia tone within a circular frame.

Bea Costa Gomes is a researcher in the Futures Team at Microsoft AI. She’s fundamentally curious about humans, and is researching Human-AI interactions with a particular interest in health. Before joining MAI, Bea was a Research Fellow and science communicator/podcast host at the Alan Turing Institute in London.

English (United States)
Your Privacy Choices Opt-Out Icon Your Privacy Choices
Consumer Health Privacy Sitemap Contact Microsoft Privacy Manage cookies Terms of use Trademarks Safety & eco Recycling About our ads