When and How to Use AI: A Framework for Reflection
This post was prepared for a guest lecture in DESINV218 at UC Berkeley. Slides for the presentation are here.
Introduction
In this post, I will discuss my approach for deciding when and how to work with generative AI tools in a research context. While I will primarily focus on the setting of academic social science research of the sort I am familiar with, my hope is that this framework might be useful for others who are trying to decide when and how to work with AI in their work (if at all), and generally build an intentional relationship with this emerging technology.
Note: this post assumes that the reader is familiar with the basics of different generative AI tools and their forms (e.g. ChatGPT, Cursor, Claude Code, etc.); useful introductions to these tools have been written elsewhere by others. The goal of this post is to reflect on when and how to use these tools.
Task Framework
One way to approach the decision about whether and how to use AI is to think in terms of tasks. In my work, a broader research project is made up of many smaller “tasks” that need to be completed. For example, tasks might include:
- developing a research idea;
- searching for related literature;
- reading related papers;
- developing a research plan;
- collecting data;
- cleaning data;
- analyzing data;
- building charts and tables;
- writing up results;
- requesting feedback;
- responding to referees
… and so on.
For each of these tasks, it may or may not make sense to work with AI when completing the task. So, given a particular task, how should we decide whether (and how) to work with AI?
Personally, I try to make this decision by considering several interlocking sub-questions, which I will discuss in the rest of this document:
If answer to this Question 1 is “yes”, we can then proceed to the following:
In other words, the questions are “can I?”, “should I?” and “how?”. In remainder of this document, I will discuss some considerations that I think about in answering each of these questions in the context of using AI in academic social science research.
Question 1: Can AI help here in principle?
Before we can decide whether and how to work with AI, we need to decide if AI can be helpful for a given task in principle. In short, is this a task where working with AI is even an option?
In some cases, the answer to this question is straightforward. For example, tasks that can directly be reduced to generating images, text, or code given some prompt can clearly be accomplished by or with the help of AI tools. This could include tasks like drafting an abstract, writing data cleaning code, or generating an illustrative figure.
In other cases, the role of AI might be more indirect. Generative AI tools can be used in a lot of different ways, and it’s worth being experimental and creative in considering potential uses. A wide range of tasks can be facilitated by generating text (or code), even when text generation is not the task per se. For example, a generative AI tool could conceivably help generate:
- Instructions on how to approach or complete a task “by hand” (without AI).
- Code to complete a task rather than executing the task directly.
- Planning or organizational support to work through a task or project.
- Suggestions of papers to read to learn more about a topic.
To give a more concrete example: perhaps you would like to conduct some face-to-face interviews with research participants. Generative AI is not great fit to execute these interviews directly, but could e.g. help generate an interview script, or make suggestions of how approach this task carefully (i.e. versions of point 1 above). There are many such possibilities depending on the task and context, and the only real way to identify them is to experiment in your particular setting to see what works. This said, I will here offer two general considerations for triaging AI opportunities based on my experience:
- “Going meta” on a task often presents opportunities for AI engagement. As alluded to above, AI might not help execute a task directly, but might help in thinking about how to approach it. (Or thinking about thinking about how to approach it etc.)
- Providing context is a key constraint. Many tasks require particular situational understanding or context about a project to be able to accomplish. So, the ability to provide your model with such context is a key constraint for its ability to be helpful with a given task at any altitude. (There is plenty of other writing about managing context, so I will not dive in too deeply here.)
A final point I will make in this section is that the answer to “Can AI help here in principle?” also clearly depends on what AI model or tool you are working with. A few points are worth mentioning about this.
First, different models and tools have different specialties which are often advertised by providers – e.g. ChatGPT Codex advertises itself as “Built for real engineering work, across multiple surfaces” – and so it’s worth knowing about these different possibilities if you are looking for opportunities to work with AI.
Second, I’ve found that personal experimenation is helpful for building intuition about the capabilities of different tools. By trying out different tasks with different models, you can build a sense of what works best for you in different settings.
Third and finally, it’s worth keeping in mind that answers to this question will evolve as AI tools and models continue developing; accordingly, continual re-evaluation of approach seems reasonable to me.
Questions 2 and 3: Should I work with AI for this task? And if so, how?
The previous section (Question 1) focused on the question of whether AI can be be helpful for a given task in principle. However, just because we can work with AI for a given task doesn’t mean that we necessarily should, for a variety of possible reasons. Hence, in this section, I will assume that the answer to question 1 is “yes”, and we now must make the decision of if and how to work with AI for a given task.
At the outset I will say that there is a lot to consider here, and I am not going to pretend to capture everything. “Should” can include a lot of things, and it’s helpful to think about what your deeper goals are in the work you are doing, what is important about it, what are your priorities, and so on. Knowing as much as possible about AI, how it works, what are its strengths/weaknesses, what are its societal risks/implications and so on, will also surely help in refining your answers.
For the purposes of this post, I am going to largely bracket some of the broader societal questions currently surrounding AI which can be treated more richly elsewhere – e.g. energy usage / environmental impact, labor market consequences, questions about intellectual property & attribution, AI safety, and so on. All I will say is that these considerations are many, and are worth engaging with in much the way that you might for any choice you make about contemporary consumption or usage. For some people, these considerations might outweigh any conceivable benefits of working with AI tools in much the way that some people choose not to eat meat, prefer to drive an electric car, or boycott social media.
In the remainder of this post, I will assume that the reader is not someone who is categorically opposed to working with AI, and will offer some narrower further considerations for thinking about whether and how to work with AI for a given task. In particular, I have personally found four questions to be helpful in making this determination:
- (i) How much do accuracy / correctness and "getting the details right" matter for this task?
- (ii) How easy is it to verify whether this task was completed accurately once it is done?
- (iii) How much more difficult is it to do this task “by hand” vs. with AI?
- (iv) What is lost/gained in doing this task “by hand” vs. “with AI” more broadly?
In the remainder of this section, I will reflect a bit futher on each of these questions, and why I think it is relevant to the decision of whether and how to work with AI for a given task.
(i) How much do accuracy / correctness and “getting the details right” matter for this task?
This question is important for working with AI because, at least in their current form, LLM tools are fundamentally fallible, sometimes in non-obvious ways. For many research tasks, a key concern is hallucination, wherein an AI tool might invent facts or citations, sometimes confidently asserting claims that are completely wrong.
But even aside from hallucination, working with AI often requires describing a broader task or goal to a model, and allowing the model to fill in the execution details (much like a manager might delegate a task to another human worker). In doing this, there is inevitably some risk of misspecification or misinterpretation. This risk becomes more of a concern for tasks where getting the all the details right is especially important.
So in short, the more that accuracy / correctness and “getting the details right” matter for a given task, the less likely I am to work with AI for that task. Or if I do work with AI, it will be in a more “hands on” way (see below).
(ii) How easy is it to verify whether this task was completed accurately once it is done?
This question is a complement to (i). If it is relatively easy to audit or verify whether a task was completed accurately, then I will be more willing to delegate that task to AI, even if accuracy is very important (since I can easily confirm whether the task was completed in the desired manner).
An important corrolary to this point is that–at least for cases where quality matters–I am unlikely to delegate a task to AI that I would not be able to complete myself in principle. If some aspect of task quality matters, I need to have enough understanding of that task quality to verify whether or not is present.
For example, if I am running a statistical analysis on some data, I am not going to delegate the analysis to AI unless I am confident (or able to become confident) that I can confirm that the statistical analysis makes sense and was done correctly.
(iii) How much more difficult is it to do this task “by hand” vs. with AI?
This is the basic cost-benefit question. Will it actually be easier to work with AI for this task, and to what extent? If a task is much easier to do with AI compared to “by hand”, I am naturally more inclined to work with AI for that task.
(iv) What is lost/gained in doing this task “by hand” vs. “with AI” more broadly?
I think about this final question in three main senses.
First, it’s worthwhile to think about opportunity cost of working with AI (or not) for a given task. Working with AI for one task might allow me to focus my cognitive effort on another, higher leverage task instead. For example, when building the front-end design for a website, allowing AI to handle the HTML/CSS implementation details might free me up to focus on the more important and foundational question of “How should this website actually be designed?”
Second, I find it helpful to think about the broader, longer-term implications – for me, for others – of working with AI for a given task.
When thinking about longer-term personal consequences, a key question is often about learning and skill development or atrophy. If I repeatedly delegate a task to AI, I am less likely to build or retain the cognitive capacity to execute that task myself. This is not inherently good or bad – it depends on the task and my priorities. Is this a skill I want to build myself, or a topic I want to learn about? Do I want to become a better writer, coder etc, or do I care more about other things?
The skill atrophy question can sometimes feel remote; after all, delegating a task to AI a single time is unlikely to undermine my capacity. But “just this once” can readily become a repeated habit, especially when using AI often offers a short-term win of completing a task quickly. A helpful thought experiment for accessing long-term consequences is: how would my life be if I always relied on AI for this type of task? How would I be if I never wrote my own emails, or wrote my own code, or did my own research? etc. As Kant shows us, a version of this question can also be used to think about bigger picture implications for the world. How would the world be if everyone always relied on AI for this type of task? Would that be a world we want to live in?
A third and final aspect of the “what is lost/gained?” question involves reflecting on what is actually important to me about a given task, and how that relates to working with AI for the task.
For example, suppose I am writing a “thank you” note or apology for a friend. For such tasks, doing them myself and in my own voice is an important part of what makes these tasks worthwhile. Even if AI could write my “thank you” more quickly and adeptly, using it would undermine an important aspect of the purpose of the task, which is me personally sitting down to express my appreciation. I feel similarly for many creative or expressive tasks. Expressing my own perspective is work that only I can do.
While many research tasks have clearer “right” and “wrong” answers, this is not exclusively the case. For example, explaining why a piece of research is important to the world is a more personal, expressive question. Of course, AI can readily provide a plausible-sounding answer to this kind of question; but it won’t be my answer – it won’t necessarily reflect what I find resonant, important, wortwhile about the research.
How should I work with AI for this task?
In most of this section, I have discussed “using AI” for a given task as though it is a black-and-white choice (“to use” or “not to use”). However, in reality, there are many different ways that I might work with AI for a given task.
For my work, the main aspect of “how to use AI” that matters to me is the extent to which I am willing to abstract myself away from the details of the execution of the task. In other words, am I wholly delegating this task to AI? Or am I working with AI in a more hands-on, interactive way? There is a spectrum of possibilities here which roughly correspond with different sorts of AI tools. Here is roughly how I think about it:
| Level of abstraction | Human/AI role | Example |
|---|---|---|
| 0. Not abstracted (Zero AI) | Human fully completes task without any AI engagement | N/A |
| 1. Not abstracted (AI advisor) | Human fully completes task, but with advice or brainstorming support from AI. No output directly generated by AI | Using ChatGPT to brainstorm an essay idea, then writing yourself |
| 2. Partially abstracted (AI assistant) | Human mainly completes task, but with assistance from AI | Writing / coding in Cursor, and periodically asking AI for help e.g. with formatting, rewording etc. |
| 3. Partially abstracted | AI mainly completes task, but with oversight, direction, feedback from human | Claude Code w/ more detailed oversight & code review; Cursor, implementing and reviewing one step at a time |
| 4. Fully abstracted | AI completes whole task, human reviews output | Claude Code w/ ignore all permissions |
The answers to the four questions (i-iv) above help me decide how to place myself on this spectrum.
For example, if (i) accuracy is very important in a given task and (ii) it’s difficult to verify accuracy, I am likely to stick closer to one of the “not abstracted” options, retaining a high amount of granular control over the task execution. Likewise, if execution details matter very little to me and it’s easy to verify accuracy, I am more willing to move towards more complete abstraction, wholly delegating the task to AI.
Evaluating Some Example Tasks
The previous section of this post provided a framework for reflecting on when and how to work with AI for a given task. In this section, I will apply this framework to several example tasks to give a sense of how I apply it in practice.
Example task: formatting a LaTeX table in a research paper.
Details: LaTeX is a typesetting language; the user writes code that looks like the image on the left below. The code is then "compiled" to generate a professional-looking table like what is on the right (example from a recent paper).
Is formatting (or reformatting) a LaTeX table a good task for AI? Let’s apply the framework from above to think about this:
So, overall, given my personal priorities, formatting a LaTeX table is a task that I am fairly comfortable delegating to AI, with the caveat that I should carefully validate for correctness afterwards.
Now, let’s consider a different task.
Example task: cleaning a raw dataset for analysis.
Details: Before any analysis can happen, a raw dataset usually needs cleaning — handling missing values, recoding and constructing variables, merging files, filtering observations, and so on. This involves many small judgment calls, each of which can be consequential for the final results. Could I hand this task off to AI?
So, at least for me, “cleaning a dataset” is something I am less willing to largely delegate to AI. If I do work with AI here it would be in far less abstracted way; for example, I might have AI propose or implement the task one step at a time, and manually check/confirm that each step makes sense before moving forward.
A few more rapid-fire examples from my perspective:
- Asking for comments / feedback on existing work; brainstorming; making suggestions. Good task. Why? Essentially, because the accuracy/quality stakes are very low. If AI gives 10 comments on my work and 9 are bad/unhelpful, I can just ignore those and address the one helpful comment.
- Generating an abstract for an already-written paper. Borderline task for reasons above.
- Finding related research papers to read on a topic. Pretty good task, again because free disposal.
- …
Again, the goal here is not to tell you that you should draw the same conclusions as me about when and how to work with AI. These conclusions depend on your own goals, priorities, and tasks.
Complicating the Task Framework
So far in this post, I have assumed that the set of tasks that make up a given project is fixed at the outset, and that our choice, consequently is whether and how to use AI for a given task.
However, this is not necessarily the case. In some cases, access to AI might present new and different ways to organize or approach a project. In other words, AI might allow us to change the set of tasks, perhaps enabling us to transform one task into a different one, or presenting the possibility of approaching a set of tasks in a new way.
Let me give you an example. In a recent academic project, I wanted to manually review a set of around 100 observations in a data set for validation purposes. The trouble was that these were complex observations, involving multiple screenshots as well as structured data, which all needed to be examined together.
Could AI reasonably play a role in this task, taking over the verification for me? Naively deploying the framework above would suggest an answer of “probably not”. After all, this is a task where accuracy is paramount, its entire value lost if delegated to AI.
However, with a bit of re-imaginging of the task, a more reasonable use of AI emerged. Using Claude Code, I was able to build a custom app interface to support the manual inspection that I planned to do; it displayed all the key information for each observation including associated screenshots and metadata, and even included LLM-generated “hints” for the manual validation task I was doing. The interface allowed me to then mark the desired validation fields, which were stored for further analysis.
Here is a screenshot of the interface:

Ultimately, the interface allowed me to complete the validation task much more quickly and accurately than I would have been able to do otherwise.
Building the interface itself was a “good” AI task because I could readily verify that the interface worked in the way that I hoped, and the details of the implementation (or understanding it) were not important to me (since I am not trying to become a software engineer).
Of course, I theoretically could have built something like this in a world without LLM tools — but the cost and complexity would have been so unreasonable given the scope of the task that I probably would never have considered it.
In short, AI here presented an alternative way to approach the broader task of “manually inspect and validate this data”. Rather than directly offload this task to AI, I could instead use AI to build the interface and facilitate its completion more readily.
Identifying these kinds of possibilities requires a bit more creative work. I’ve found that this creativity is facilitated through experimentation, observation, and familiarity with the range of possible ways of engaging with LLM tools.
What does this look like for me?
To conclude, I want to share a few final thoughts about where the above considerations have landed me in terms of my practical AI worklow as an academic social science researcher.
Personally, I find that most of my research tasks live somewhere between 0 and 3 on the AI abstraction spectrum outlined above. I will rarely fully delegate a research task to AI, mainly because accuracy/quality are very important for my work, and I want to check and keep a close eye on any research steps I am taking.
Because of this, for most of my personal research work, my preferred interface is Cursor for a few main reasons.
First, working within Cursor, I can readily incorporate all the context that I need. I typically try to organize my projects so that there is one high-level folder that includes various project aspects (e.g. data, code, results, write-ups, references, planning materials etc.); I can then open Cursor at this altitude for shared context, even if I am working on a specific part of project only.
Second, Cursor is very flexible across different types of AI uses. Sometimes, I just want to ask a question with relevant context, sometimes I want to brainstorm, sometimes I want to get feedback or make changes, sometimes I want to build a whole thing in more agentic mode. Cursor allows me to do all of these with relevant context.
Third, I find the Cursor UI very useful; it’s nice to be able to refer to specific lines or files, and to generate new documents in the same space etc and see clear diffs of what was added or removed.
Finally, since Cursor is built on VSCode, it fits well with my existing workflows (which previously relied on VSCode), and can play nicely with other helpful VSCode plugins, e.g. for markdown or LaTeX formatting.
In addition to Cursor, I’ll tend to reach for Claude Code if I want to do a more fully-delegated AI task, where I am less worried about the details.
I will conclude by reiterating a final time: this is just the approach that works for me, based on my above thinking, priorities, and the types of tasks that I tend to do. You might have different priorities or different tasks which imply I different set of choices about whether and how to work with AI, if at all.