In this episode, I talk with Ning Leng, Ph.D., Director II, Data and AI Acceleration Group, Data & Statistical Sciences at AbbVie, about the growing role of the R Consortium in the pharmaceutical industry.

Ning brings extensive experience in statistics, computational genomics, open-source technology, and the adoption of R across the pharmaceutical industry. Before joining AbbVie, she spent 10 years at Roche Genentech, where she helped drive the adoption of R, cloud technologies, Git, and Shiny.

We discuss how the R Consortium creates a platform for statisticians, programmers, pharmaceutical companies, and regulators to collaborate on practical challenges—and how that collaboration is changing the way we approach regulatory submissions.

Why you should listen

If you work in statistics, data science, statistical programming, or clinical development, this episode will give you a practical look at how the industry is moving toward more modern and collaborative approaches.

You’ll learn:

  • How the R Consortium is helping pharmaceutical companies work together and collaborate with the FDA.
  • What the R Consortium’s regulatory submission pilots can teach you about using R in real-world submissions.
  • How open-source collaboration can influence regulatory guidance and industry practices.
  • How tools like R, Shiny, containers, and WebAssembly could make regulatory submissions more interactive and reproducible.
  • How the community is beginning to explore generative AI for clinical trials and statistical programming.
  • Why benchmark datasets, challenging test cases, and quality control will be critical for using AI responsibly.
  • How you can explore the R Consortium’s public resources and get involved in its work.
  • Whether you’re already using R or simply want to understand where statistical programming and regulatory submissions are heading, this conversation with Ning Leng will give you valuable insights into the future of our field.

Episode highlights with timestamps

  • 01:30 — Meet Ning Leng and her journey into R
    Ning introduces her background in statistics and computational genomics and explains how she became involved in R adoption, cloud migration, Git, and Shiny at Roche Genentech.
  • 03:57 — Why R wasn’t being used for regulatory submissions
    Ning explains the misconception that the FDA did not accept R and identifies the real challenge: the industry lacked practical examples showing how to submit R-based materials.
  • 05:46 — What is the R Consortium?
    We discuss the R Consortium’s role as a nonprofit organization that promotes good practices and the use of R across industries.
  • 06:16 — How the R Consortium collaborates with the FDA
    Ning explains the submission working group, validation hub, and how the Consortium provides a platform for collaboration between pharmaceutical companies and the FDA.
  • 08:13 — R Consortium FDA submission pilots
    Ning walks through the different pilots, from submitting TLGs to incorporating Shiny, ADaM code, containers, WebAssembly, and alternative data formats.
  • 10:29 — How companies are using the pilot submissions
    We explore how pharmaceutical companies use the public R Consortium pilots as practical templates when preparing their own R-based regulatory submissions.
  • 11:24 — How collaboration influenced FDA guidance
    Ning explains how lessons from the pilots helped clarify FDA guidance around file formats, including .r and .zip files.
  • 12:17 — Moving beyond PDF-based electronic submissions
    We discuss the opportunity to use interactive graphics, HTML, Shiny, and other modern technologies to make electronic regulatory submissions more useful and traceable.
  • 14:14 — Generative AI enters the picture
    Ning explains how the R Consortium is beginning to explore generative AI for clinical trial reporting, trial design, and statistical programming.
  • 15:12 — Building AI skills and benchmark test cases
    We discuss open-source AI skills, benchmark datasets, and the importance of testing AI against difficult and unusual corner cases.
  • 16:28 — How you can get involved with the R Consortium
    Ning shares practical ways to explore the community, including its public working-group materials, meeting minutes, recordings, and Slack channel.

Links and Resources:

🔗 R Consortium: Learn more about the R Consortium and its work to promote the use of R and good practices across industries: https://r-consortium.org/

🔗 R Consortium R Submissions Working Group: Explore the working group’s projects, meeting materials, submission pilots, and opportunities to get involved: https://rconsortium.github.io/submissions-wg/

🔗 R Submissions Working Group — Pilot Projects: Learn about the different FDA submission pilots and how the community is exploring R for regulatory submissions: https://rconsortium.github.io/submissions-wg/pilot_background.html

🔗 R Consortium 2026 Plans and 2025 Success: Read about the latest work from the R Submissions Working Group, including the development of Pilots 6 and 7: https://r-consortium.org/posts/submissions-wg-2026/

🔗 Pilot 4 — WebAssembly and Containers: Learn how the R Consortium explored submitting a Shiny application using WebAssembly and containers for FDA review: https://r-consortium.org/posts/using-r-to-submit-research-to-the-fda-pilot-4-successfully-submitted/

🔗 R Consortium Working Groups: Browse the R Consortium’s different working groups and projects: https://r-consortium.org/all-projects/isc-working-groups.html

🔗 The Effective Statistician Academy – I offer free and premium resources to help you become a more effective statistician.

🔗 My New Book: How to Be an Effective Statistician – Volume 1 – It’s packed with insights to help statisticians, data scientists, and quantitative professionals excel as leaders, collaborators, and change-makers in healthcare and medicine.

Join the Conversation:
Did you find this episode helpful? Share it with your colleagues and let me know your thoughts! Connect with me on LinkedIn and be part of the discussion.

Subscribe & Stay Updated:
Never miss an episode! Subscribe to The Effective Statistician on your favorite podcast platform and continue growing your influence as a statistician.

Never miss an episode!

Join thousends of your peers and subscribe to get our latest updates by email!

Get the shownotes of our podcast episodes plus tips and tricks to increase your impact at work to boost your career!

We won’t send you spam. Unsubscribe at any time. Powered by Kit

Learn on demand

Click on the button to see our Teachble Inc. cources.

Load content

Ning Leng, PhD

Director II, Data and AI Acceleration Group, Data & Statistical Sciences, at AbbVie Inc.

Ning Leng recently joined AbbVie’s Data & AI Acceleration group lead end-to-end in-silico simulation efforts that optimize clinical trial decisions and advancement. Prior to Abbvie, Ning served as the Global Head of Data Science Acceleration group at Roche-Genentech, where she led large-scale modernization efforts across clinical reporting—spanning end-to-end R-based regulatory filing, Shiny-enabled trial analysis dashboards, and GenAI productivity tools. Earlier, Ning was a statistician supporting early-phase oncology development and biomarker discovery.

Ning is also a cross-industry advocate for clinical data science modernization. She co-founded and co-led the R Consortium Submissions Working Group, which has partnered with the FDA across five pilots to demonstrate the feasibility of modern technologies for regulatory review and submission, including R, Shiny, containers, and WebAssembly. Ning earned her PhD in Statistics from the University of Wisconsin–Madison and her bachelor degree in Information and Computing Science from the Beijing Institute for Technology. 


Transcript

00:00
You are listening to the Effective Statistician Podcast, the weekly podcast with Alexander Schacht and Benjamin Piske designed to help you reach your potential, great science and serve patients while having a great work-life balance.

00:21
In addition to our premium courses on the Effective Statistician Academy, we also have lots of free resources for you across all kind of different topics within that Academy. Head over to theeffectivestatistician.com and find the Academy and much more for you to become an effective statistician.

00:49
I’m producing this podcast in association with PSI, a community dedicated to leading and promoting the use of statistics within the healthcare industry for the benefit of patients. Join PSI today to further develop your statistical capabilities with access to the ever-growing video-on-demand content library, free registration to all PSI webinars, and much,

01:13
head over to the PSI website at psireb.org to learn more about PSI activities and become a PSI member today.

01:30
Welcome to another episode of the Effective Statistician. And today I’m super happy to talk about a topic that was on my radar for quite some time, but I never found the right person to actually talk about it. And therefore I’m super happy that Ning is today on the show and we can speak about the R-Consortium. So welcome to the episode and see the podcasting, maybe you can start by

02:00
introducing yourself. great. Thank you so much for having me today. So my name is Ning and right now I’m at, at WE. I recently joined WE about three months ago. And before that I spent 10 years at Roche Genentech. I was trained as a statistician. So I graduated from University of Wisconsin, Madison with a focus on computational genomics. After that I stayed in academia a little bit. At Roche I started as a statistician for the first five, six years.

02:29
And then kind of like on the side, because I don’t really know SaaS coming to the industry. So myself and also several friends within Roche, we actually started a user community at that time. Kind of teaching people how to like write R packages, how to write shiny apps and how to use Git. then my fifth or sixth years learning at Roche GenTech, basically our leadership team decided to have this strategy of adopting open source languages and modernized technology.

02:58
So I got moved to an innovation team. I got the opportunity to lead an innovation team to drive adoption of R, to drive the cloud migration and also the kind of Git adoption and Shiny adoption, et cetera. So that was kind of how I got involved in this whole very exciting R transformation in the cross industry world and also within Roche Genentech at that time. think there’s a couple of companies that have invested quite a lot into R.

03:28
Roche was definitely one of the early movers. Many more companies that go into that direction, but I think none of the companies have as radically transformed as Roche did, at least looking from the outside. So the R Consortium, did you personally got involved in that? How did that happen? So basically it kind of was triggered by some internal discussion about why are we not using R for regulatory submission.

03:57
Actually, adoption of R, I think within Roche, we have been discussing that for more than 10 years. What I heard was even in 2010 years time, people actually started talking about R and such. As a group, we try to figure out why the industry is not using R for regulatory submission. In the beginning, people were like, oh, because FDA doesn’t accept that.

04:20
But if you do a little bit more research, you realize that since 2015, think FDA issued a statement saying that FDA doesn’t require any specific programming language and you just need to make sure that your program is traceable and have high quality. And they have been repeating the same message during many conferences over there. So basically FDA does assess R and also for the kind of younger generation reviewers, we learned that they come up the school with R skill, Python skill.

04:49
And they already use R in their day-to-day job anyways. And we realized the bottleneck is really that there is no good example showing people how to use R for regulatory submission. Like how to put your full materials in a certain format through the ECTD submission portal so that the FDA reviewers can retrieve the information they needed during review. So with that, we kind of decided to have a working group working together with FDA.

05:18
to do some kind of mock submissions to FDA and put everything in the public so that for companies who want to do our submission, they can use those mock submissions as a template to show them how to do certain things. This is so cool. And I remember when I left university in 2002, so that is 24 years ago, I was already using ARND, although I was more for SARS.

05:46
trained person, all the new younger people at that time were already using R and all the packages were already done in R. And so it took quite some time to get here. But I think the R Consortium actually played a very, very critical role in that. What’s the role of R Consortium in the whole R space? That’s a very good question. So I think R Consortium as an organization, basically it is not far more specific.

06:16
So it is a non-profit organization trying to promote the good practices and usage of R across different industries. During the recent years, I see pharma definitely being one of the most active space with the most number of working groups for R Consortium because our recent adoption of R and open source language. yeah, under R Consortium, there are several working groups. think the most active two or three are like the submission working group where we

06:46
work together with FDA on a number of pilot submissions, trying to showcase how to submit certain component to FDA. And there is also our validation hub where people found the materials really, really useful. They try to come up with white papers to document what does validation mean for open source language. And also they open source some of the tools over there in terms of quality metrics, measurement, et cetera.

07:14
And I also know there are several conferences, R-related conferences that our consortium is supporting. Yeah. So I definitely feel like our consortium as a non-for-profit organization provided this very nice platform for us to collaborate across the companies and also with FDA. So basically the R Consortium is the organizational backbone of the whole R community across different industries. And it’s great to hear that the pharma industry could

07:43
on board and now is very active and I’m pretty sure many other industries can benefit from it and vice versa. So you already mentioned the submission case studies. So what do you see kind of how are these leveraged by the different companies? Is this embraced or is that still some kind of outlier here and there or what’s the status of our based submissions?

08:13
Right. Yeah. Maybe I can first give an overview about the pilot that we’re running. And then I can also share that like how the group practice evolve over time, like based on our learnings, et cetera. So for Arc of Social Mission Working Group, right now we are at pilots six and seven. We already finished four pilots. There are four pilots that already have FDA response letter.

08:38
Recently, the number four was completed. If people are interested, the letter is on our website right now. So the first one was very simple. Panel one was very simple. Basically, we tried to submit four TLGs using the CDISPILOT1, the simulated data to submit four TLGs. We added an ECD portal and then to make sure the FDA reviewers were able to use the exact same package, same environment to reproduce these four TLGs.

09:06
In the for panel number two, we add a Shiny component. So basically we believe that right now with the trials being more more complicated and more and endpoints, more and more complex measurements, etc. like Shiny could be a really, really nice supplementary tool to accelerate review. So for panel two, we success a Shiny app to FDA. And in panel three, we submitted in addition to TLF code, we also submitted Adam code.

09:35
So that to enable the FDA reviewers to reproduce the codes to generate the atom as well. In panel four, the recent one was a very bold one. Basically we submitted a Shiny app through container in web assembly. So the FDA reviewer was able to review like reproduce the environment using Docker in web assembly. But there were also some IT challenges that’s kind of like highlighted in the response letter.

10:02
In pilot five, we are trying to submitting the dataset in JSON format instead of XPT format so that people don’t need to do this like artificial conversion from a JSON data file to XPT data file over there. then pilot six and seven are kind of like our playground right now to try out different AI solutions. Pilot six is trying to use AI to generate the element LF. And pilot seven is trying to

10:29
like create benchmark data set and benchmark test cases for AI tools and also kind of traditional open source tools. So that’s where we are right now. Back to your question about how does this impact the companies like real submission. What I heard was definitely that I heard from multiple companies like Novo and DNJ that when they did their first R submission, they closely studied our pilot submission cases.

10:58
and try to mimic what we did and try to avoid the bottleneck or the caveats we encountered at that time. Roached the same that we try to follow the open pilots learnings over there. And the other thing I want to call out is that also because of those pilot submissions, we were able to influence FDA to update some of their guidance or documents. For example, like in the past in their

11:24
file format guidance, they said like no zip file and they didn’t specify too much about like the files with .r like suffix. So we were not sure whether we can submit a proprietary R package because oftentimes that come with a zip file, the suffix. And also we wasn’t sure how to submit a .r file. So in the very beginning, even artificially

11:48
converted the R file to .txt file. Last year, actually FDA was able to update their file format guidance, specifically saying that, we allow .zip file for our package submission, and we allow for .r, .etc, like, e-Base affixes around the R ecosystem. So with that, actually, that streamlines a lot of the operational side of the submission as well. That is so cool. This overall collaboration and the influence it has on.

12:17
various companies, but also on the FDA and it shows how people, if they have a common goal, can work very, very effectively together. I also think that that will overall help quite a lot to make process better, like the shiny thing. Have you done anything that is more HTML based? So some kind of a markdown so that you don’t need a container or something like this, but just an HTML file?

12:46
to have some kind of interactive graphics. We haven’t submitted an HTML, but I feel like the web assembly portion of the Pilot 4 is kind of toward that direction. Basically for the web assembly approach, a reviewer will be able to open the Shiny App in their own browser without too much of environment setup. So that one actually went through quite smoothly, I would say. Because I think that is definitely part of the future,

13:16
We don’t submit documents where the only difference to the pre-electronic submission is that it’s now a PDF or document instead of an actual binder, physical binder. the biggest advantage is having hyperlinks, but now actually embracing what’s possible with electronic submissions, having interactive things, having much more kind of traceability.

13:44
All these kinds of different things is there’s so much opportunity in the future. And I’m really curious what will happen in terms of the generative AI. I’m really looking forward to that because I would personally have lots of questions around that. Because I’ve used or I tested AI for producing text around tables and things like this. And it’s not as, as I would like it to be still, but what’s your perception up to now?

14:14
Yeah, that’s exactly kind of like our motivation of like moving to pilot seven, where I mentioned that we started to think about what does benchmark mean for the clinical trial reporting and also for trial design, basically the biostandard and programming related tasks. So for pilot seven, actually what we’re doing right now is that we ask the community to contribute their AI skills. As you probably know that basically.

14:42
Skill is a very powerful approach to provide SOP to an AI algorithm to tell them when to use certain packages, when to use certain tools and how to QC before return a result to me, et cetera. So we asked the community to open source several skills they’re writing. So right now we have three set of skills. One is about sequential design. One is about trial simulation and one is about

15:12
item generation. And after open sourcing those skills, right now we are also asking the community to contribute benchmark test cases. So we ask people to provide us the craziest test case or the most weird corner case they encountered. And we try to improve those AI tools based on those corner cases.

15:34
We’re still at a very early stage. think we only started probably like two months or a little bit over a month ago, but there is a very strong contribution over there. I saw that there are like over 50 stars in the repository and we have over a hundred issues now. So hopefully with that this approach, we will come up with some high quality skills so that it can help us to manage the.

15:59
kind of like our certainty around generative AI a little bit better and then to kind of really unleash that in clinical trial setting. I have the feeling there is another episode in that. Thanks so much Ning for this awesome discussion and getting us information about the role of the R Consortium, what’s overall heading going on in the community, how much that is impacting our

16:28
industry, our collaboration with the FDA, and also to have a look out what further things will come in the future. Thanks so much, Ning. One final thing, if someone is more new in the R Consortium area, what would be good things to check out and learn about? Yeah, that’s a good question. So for our consortium, one thing I’m really proud about is everything is public.

16:58
So for anyone who is interested in our working group or any of the projects, all the meeting minutes are on our website and all the recordings you can also find on our website. And also, everyone is welcome to join our Slack channel. So I think for people who are exploring, it will be nice to maybe go through some of the meeting minutes and also go through the Slack discussions and identify which specific area they want to get involved in.

17:25
We’ll put that all into the show note, the different links to all of that. Of course, the link to Ning’s address and page so that you can follow her there as well. Thanks so much Ning. Sounds great. This show was created in association with PSI. Thanks to Rain and her team FEVS. Well, put the show in the background and thank you for listening. Reach your potential, leak great silence and serve patients. Just be an effective statistician.

17:59
This show was created in association with PSN. Thanks to Reign and her team at VVS for help with the show in the background and thank you for listening. Read your potential, leak great science and serve patients. Just be an effective statistician.

Join The Effective Statistician LinkedIn group

I want to help the community of statisticians, data scientists, programmers and other quantitative scientists to be more influential, innovative, and effective. I believe that as a community we can help our research, our regulatory and payer systems, and ultimately physicians and patients take better decisions based on better evidence.

I work to achieve a future in which everyone can access the right evidence in the right format at the right time to make sound decisions.

When my kids are sick, I want to have good evidence to discuss with the physician about the different therapy choices.

When my mother is sick, I want her to understand the evidence and being able to understand it.

When I get sick, I want to find evidence that I can trust and that helps me to have meaningful discussions with my healthcare professionals.

I want to live in a world, where the media reports correctly about medical evidence and in which society distinguishes between fake evidence and real evidence.

Let’s work together to achieve this.