Research on fan fiction & large language model training data!
Do you write fan fiction? Do you have feelings about things like ChatGPT or Bard? I'm Katy Gero and as a Postdoctoral Fellow at Harvard I am running a study on how writers feel about their work being used to train large language models. In addition to writers publishing with big and small presses, I'm interested in interviewing fan fiction writers to understand their perspective on this issue.
The goal of this research project is to better understand the concerns of different writers, and ultimately move toward responsible, ethical, and consensual data collection practices in the future. This project does not presume writers do or should want their writing to be used as training data, but rather seeks to understand if there exist conditions under which writers would want their writing used---and if those conditions can feasibly be met.
By being interviewed, this academic work can better represent your opinion, and has the potential to shift the landscape of how data collection and licensing is done in the future.
The interview would take about an hour and be conducted on Zoom. In accordance with the ethical research practices at Harvard University, audio recordings of interviews are stored for no more than a year, shared only with researchers on the project, and any results from the interviews (such as themes or quotes) are fully anonymized before being shared publicly (e.g. in an academic publication).
You can find out about my research on AI and writing more generally at my website, www.katygero.com.
If you are interested, please fill out this super quick screening survey. Then I can be in touch about scheduling the interview: https://forms.gle/VMvesuWw9QCuRiSP8
Whether you participate or not, please consider sharing this survey with your social networks!
If you have any questions about the project or otherwise, please don't hesitate to email me: [email protected]
Much thanks!















