No history yet

I want to learn how to create good quality research data management plan in the field of theoretical linguistics.

I would love to help you design a solid, practical Data Management Plan (DMP) for your theoretical linguistics work. Because linguistics data can range from delicate native-speaker audio recordings to complex syntactic treebanks, tailoring the plan to your specific setup is key.

To help me customize this guide to your exact needs, could you tell me a bit about your current research project and what stage you are at? For example, are you planning a new fieldwork study, preparing a grant proposal, or trying to organize a large pile of existing data?

I am planning a brand new fieldwork project.

A brand new fieldwork project is the perfect time to build a (DMP). In fieldwork, you are dealing with raw, irreplaceable human interactions—audio recordings, translated wordlists, and video files—that cannot simply be re-run in a lab if something goes wrong.

fieldwork

noun

The collection of primary linguistic data through direct interaction and observation of native speakers in their communities.

Before we dive into the specific structural parts of a DMP, it helps to understand why we write them. In theoretical linguistics, a good plan protects your data from physical loss, makes your analysis reproducible for other scholars, and ensures you respect the intellectual and personal rights of the community you are working with.

Continue to the core components of a standard DMP

Let's break down the blueprint. A standard Data Management Plan is built around six core components that work together to protect your research from start to finish.

The Six Pillars of a DMP

To make this highly practical for your fieldwork, let's map out exactly what these six sections look like. Think of them as the sequential chapters of your project's data biography:
  1. Administrative Information: The basic metadata of your project—who you are, what the project is, who is funding it, and who is responsible for managing the data.
  2. Methods: A clear explanation of what kinds of data you will collect (e.g., audio, syntactic treebanks) and how you will collect them (e.g., Zoom H4n recorders, specialized elicitation software).
  3. Documentation and Metadata: This is where you explain how someone else (or future you!) will understand your data. It covers your file-naming conventions, folder structures, and the descriptive schemas you'll use.
  4. Storage and Security: Your immediate safety net. This section outlines your daily backup routines, password protections, and how you plan to keep physical and digital files safe while you are traveling.
  5. Ethics and Legal Compliance: Crucial for linguistics. Here, you address how you will obtain informed consent, how you will anonymize sensitive participant identities, and who legally owns the resulting data.
  6. Sharing and Archiving: The final destination. This defines how, where, and when you will make your data available to the broader scientific community for long-term preservation.

Since you are preparing for a brand new fieldwork project, getting these six sections in order before you pack your bags will save you hundreds of hours of headache down the road.

Make me a worksheet to sketch out these six areas

I would love to help you sketch this out! Having a structured space to brainstorm your fieldwork parameters is the absolute best way to ensure nothing gets missed before you head into the field.

SHEET
Worksheet

I've set up a workspace where you can start detailing your fieldwork plans step-by-step.

Now, let's zoom in on the first real technical milestone of your plan: Identifying Linguistic Data Types. In the field, you won't just be collecting sound; you'll be creating a web of interconnected digital files.

To tailor our next step, what specific language or language family are you planning to study, and what kind of linguistic phenomena (for example, sentence structures, speech sounds, or vocabulary lists) are you primary focusing on?

I'm planning to gather data on sentential negation in Slavic languages, with Miestamo's and de Clercq as my main sources.

Studying sentential negation in Slavic languages through the lenses of Miestamo and de Clercq is a fantastic foundation for a fieldwork project. Because negation is a core syntactic and semantic phenomenon, your data will span multiple layers of analysis, from raw speaker utterances to deeply annotated structural representations.

The Stack of Negation Data

For a fieldwork project looking at how Slavic languages express negation, your raw and processed data types will likely stack up like this:

  • Audio Recordings: High-quality, uncompressed audio files (like WAV format) containing elicitation sessions, grammatical judgment tasks, or spontaneous narrative interviews.
  • Transcription Files: Written representations of those recordings. This could be orthographic (using the language's standard alphabet) or phonetic transcriptions using the (IPA).
  • Annotated Corpora & Treebanks: Structural software files (such as XML, JSON, or software-specific formats from tools like ELAN or Praat) where you align the audio with syntactic trees, parts of speech, and specific tags identifying the type of negation.

For example, when looking at symmetric versus asymmetric negation under Miestamo's typology, your metadata and annotations will need to clearly mark whether the negative sentence structurally differs from its affirmative counterpart beyond just adding a negative marker. Having structured file formats makes these grammatical comparisons queryable later.