Autonomy, sensitive data and sustainability in research settings
September 7, 2026
María Cristina Nanton
Developer
Health Ministry @ Buenos Aires | The Global Health Network
Data science in the public sector
Data-driven projects in a global health research network
There’s a new kind of coding I call “vibe coding”, where you fully give in to the vibes, embrace exponentials, and forget that the code even exists. It’s possible because the LLMs (e.g. Cursor Composer w Sonnet) are getting too good. […]
I ask for the dumbest things like “decrease the padding on the sidebar by half” because I’m too lazy to find it. I “Accept All” always, I don’t read the diffs anymore. When I get error messages I just copy paste them in with no comment, usually that fixes it. […]
I’m building a project or webapp, but it’s not really coding - I just see stuff, say stuff, run stuff, and copy paste stuff, and it mostly works.
— Andrej Karpathy, Twitter, February 2025 (emphasis mine)
Generative AI platforms that give access to advanced language models
Claude - Anthropic
(Claude Opus 5, Sonnet 5, Fable 5.1)
ChatGPT - OpenAI
(GPT-5.5 Instant, GPT-5.6 Sol)
Copilot - Microsoft
(GPT-5.6, Claude Opus 5)
Conversational interface + specific ways to integrate them into code-based projects
Claude Code | Codex | GitHub Copilot | Etc
And also open models: downloadable models you can run on your own infrastructure (a laptop, an institute server, an already-approved cloud)
Who publishes them
What you need to run them
How they run
What you gain
If it runs on your own infrastructure, the data never leaves the institution.
Products such as
Some examples I’ve seen
Some questions that often come up:
The regulatory floor
Your own floor — three questions before you start
Old problems, amplified: a Google Doc stays where you left it; a prompt can end up as training data, and an app generated with Lovable can publish its database to the internet.
Anthropic, AI for Science program — thematic call (Jul 2026)
Julien Gagneur (Technical University of Munich)
Shadow AI is the unsanctioned use of any artificial intelligence (AI) tool or application by employees or end users without the formal approval or oversight of the information technology (IT) department.
— IBM, “What is shadow AI?” (2024)
New risks: what data gets sent to commercial platforms and models, both voluntarily and involuntarily, and how a model’s outputs get absorbed into institutional assets.
Nobody meant to leak anything: each of them was solving a work problem with the tool at hand
What data?
Confidential
.env file, in the code, in a screenshotInternal
Through where?
The prompt: “let me paste the table so you can help me”
The shared folder: an AI with access to a folder reads everything in it
The terms of service: the provider stores inputs and outputs and, depending on the plan, trains on them
What the AI generates: apps with a public URL and a database with no access rules; repos with credentials inside
Integrations: connectors to Drive, email or calendar; browser extensions
Any digital asset that affects people, projects or products in a research setting calls for assessing:
But also..
By all means, prompt AI. But don’t just relay the output. Read it, understand it, validate it, and then write a response in your own words (a decent certificate that you’ve done the prior steps). Making that effort is value you can add.
— Niklas Gruhn, “Don’t be a meat proxy”, August 3, 2026
On this, 3 key questions
What value are we expected to bring
to the projects we work on?
What is worth investing time in?
Where data is stored is a decision
Is it your decision?
Thank you!
María Cristina Nanton
mcnanton@gmail.com
BIOMED Scientific Seminar · Versión en español