Academic preprints and author contribution notes are invaluable for sharing work early and crediting collaborators. They can also reveal more about you than you intended: personal email addresses, phone numbers, home or lab locations, funding identifiers, and even signatures embedded in PDF metadata. If you’re applying for jobs, changing institutions, or tightening your digital footprint, those details can persist across mirrors, indexes, and citation databases. This guide shows you how to find what’s exposed and remove or minimize it—while protecting your scholarly record.
What Personal Details Commonly Leak on Preprint and Author Pages?
Personal information can appear in several places—sometimes beyond the main PDF:
- Title page and footers: personal email, phone number, mailing address, department location.
- Author contribution notes: full names with middle initials, lab locations, ORCID iDs, direct emails for correspondence.
- Supplementary materials: raw data files with embedded metadata (author, device, GPS coordinates).
- PDF metadata: “Author,” “Producer,” and “Keywords” fields may contain names, emails, and internal notes.
- Preprint server profile pages: contact details and affiliations sync from user accounts.
- Repository or code links: Git commits with email addresses, issue trackers with contact info.
Before You Start: Quick Triage
Make a plan to avoid creating a bigger digital trail while fixing the problem:
- List every location where the paper appears: the original preprint server, institutional repository, personal site, lab site, GitHub, ResearchGate/Academia, Google Scholar profile, and any mirrors or aggregators.
- Identify which details to remove or reduce: email, phone, exact office location, ORCID visibility, middle names, personal domains.
- Create a redacted version of the paper: move contact to a role-based or departmental address, remove phones, and review footers/headers.
- Sanitize files: scrub PDF and supplementary-file metadata before re-uploading.
How to Sanitize Your Files
Strip Personal Info from the Manuscript
- Title page: Replace personal email with a role-based or institution-managed address (e.g., corresponding@dept.edu). Remove phone numbers and precise room numbers.
- Author contributions: Keep contribution statements but remove personal emails; list ORCID iDs judiciously or set them to limited visibility. Avoid naming exact lab locations if not required.
- Acknowledgments: Thank institutions and funders without listing personal contact details.
- Affiliations: Use institution and city, not street addresses.
Clean PDF and Supplementary Metadata
- PDF metadata: In your PDF editor, remove values in Author, Subject, Keywords, and Producer fields that contain personal data. If using LaTeX, avoid packages that auto-insert full names or emails into metadata.
- Image and dataset EXIF/XMP: Strip EXIF from images; remove author tags and GPS coordinates. For CSV/Excel, check document properties for author names or emails.
- Code repos: Amend commits sparingly; better, set a privacy-friendly noreply email moving forward. Remove emails from README files and issue templates.
Platform-Specific Removal and Update Paths
Every server has its own policy. Most allow metadata edits and new versions, but older versions may remain visible or accessible via version history, DOIs, or mirrors. Here’s how to approach the common platforms.
arXiv
- Update preferred contact: Change your arXiv account email to a role-based or long-term professional address.
- Upload a new version (v2+): Submit a revised PDF with redacted contact details and sanitized metadata. Explain privacy reasons succinctly in the submission comments.
- Request selective redactions: If a prior version exposes sensitive info (e.g., personal phone), open a help ticket requesting removal of the exposed line in the abstract page or to hide email from the abstract listing. arXiv rarely removes versions but may address sensitive exposures.
- Affiliation display: Ensure affiliation in the arXiv metadata omits room numbers or home addresses.
bioRxiv/medRxiv
- Contact the support desk with the DOI/identifier and explicitly list exposed data (email, phone, address). Request removal from the abstract page and replacement of the PDF with a clean version.
- Upload a corrected version following their revision process. Ask if the earlier PDF can be suppressed or have the landing-page email hidden.
- Check author notes: Some templates auto-import corresponding author email to the landing page—request a display change if needed.
SSRN
- Edit paper details to remove personal contact and use a professional or role-based address.
- Replace the manuscript with a revised version. SSRN may retain history; request that specific exposed details be removed from the abstract page and legacy files if feasible.
- Author page: Remove phone numbers and personal websites you no longer control.
HAL, OSF, institutional repositories
- Edit metadata in the repository record to remove personal emails or room numbers.
- Upload a new file and ask the repository manager to suppress or overwrite older files that contain personal details, if policy allows.
- Persistent identifiers: If the item has a DOI/handle, request landing-page edits even if the file stays.
ResearchGate and Academia.edu
- Profile privacy: Hide your email, phone, and personal links. Restrict who can message you.
- Paper uploads: Replace any uploaded PDFs that include personal details and delete outdated versions where permitted.
- Mirrored metadata: If the platform auto-creates a record from elsewhere, contact support to correct author info or remove personal data.
ORCID and Profiles That Feed Preprint Pages
- ORCID visibility controls: Set email to private or trusted-parties only. Review “Other IDs,” employment, education, and works for location or contact details.
- Publisher/aggregator sync: Some sites pull data from ORCID or Crossref. Update ORCID first, then request a refresh on downstream sites.
Contact Templates for Support Requests
When writing to support, be concise and specific. Include relevant identifiers and an explicit privacy rationale. Example:
Subject: Request to remove personal contact details from [Identifier/DOI]
Hello [Support Team],
I’m the corresponding author of “[Title],” [ID/DOI]. The abstract page and/or PDF includes my personal [email/phone/address], which I no longer wish to display. I have uploaded a revised file without these details. Could you please: (1) remove/hide the personal [detail] from the landing page metadata, and (2) suppress or replace access to prior versions that expose this information, if policy allows? I appreciate any guidance on the appropriate process.
Thank you,
[Name], [ORCID (optional)], [Institution]
What If the Site Won’t Remove Old Versions?
Some repositories keep immutable version histories. If a platform cannot delete an older version, reduce exposure and context:
- Ensure the current version is clean and set as the default or “latest” so it’s what readers see first.
- Request landing-page redaction so sensitive data isn’t displayed outside the PDF.
- Ask for robots directives to limit indexing of old versions, when policy permits.
- Update downstream records (Google Scholar, ResearchGate, institutional pages) to point to the clean version.
- Replace personal email with a long-lived role account so any remaining exposure is lower risk.
Reduce Rediscovery Through Indexers and Mirrors
Even after cleaning the source, old data may persist in caches or mirrors.
- Search operators: Query your name/title with operators like site:arxiv.org, site:ssrn.com, or filetype:pdf to find stray copies.
- Contact mirrors and ask them to pull the updated version or to remove the landing-page email from their abstracts.
- Web caches: Request cache refresh from search engines after the source page is updated.
- Institutional pages: Ask your former department to remove author bio pages that include personal emails or phone numbers linked to the paper.
Privacy-Smart Author Contribution Notes
You can maintain transparency without sharing personal information:
- Use contributor roles (CRediT taxonomy) without connecting them to personal emails.
- Centralize correspondence to a non-personal inbox or a journal message system.
- List affiliations at a high level (institution, city, country) and avoid office numbers or lab suites.
- Disclose funding properly without adding private contact lines.
Common Mistakes to Avoid
- Forgetting PDF metadata: Even with a clean title page, metadata can leak names and emails.
- Leaving emails in supplementary code/data: Notebooks, CSV “author” fields, and git history can reveal contact info.
- Posting redacted and unredacted versions simultaneously: Remove the old file first when possible, then upload the redacted one.
- Using personal domains that you might abandon; links can be captured in archives you can’t control.
Documenting the Fix
Keep a record in case you later need to demonstrate due diligence:
- Screenshot the before/after of landing pages and metadata.
- Save correspondence with support teams and repository managers.
- Note dates of version updates and cache invalidations.
- Monitor search results monthly for a few cycles to confirm the old data is gone.
Ongoing Monitoring and Identity Protection
After cleanup, it helps to monitor for unexpected exposure and identity risks tied to your name and email. Academic emails often anchor financial and professional accounts, making them targets for phishing and account takeover during job or affiliation changes. If you want continuous oversight for identity and credit risks, consider a reputable monitoring service that alerts you to new-credit inquiries, account changes, or data-leak indicators. A practical option is outlined here: SmartCredit for privacy, credit monitoring, and identity protection. Use monitoring as a complement to—never a substitute for—removing exposed details at the source.
Checklist: Fast Path to Redaction
- Identify all locations (preprint server, repositories, profiles, mirrors).
- Create a redacted manuscript and sanitize metadata and supplements.
- Upload a new version and request landing-page redactions.
- Ask support to suppress legacy files or limit indexing where allowed.
- Update ORCID visibility and profile pages that feed metadata.
- Replace personal contact with role-based addresses.
- Search and clean mirrors; request cache refreshes.
- Document actions and monitor for reappearance.
Frequently Asked Questions
Will updating the PDF remove my personal email from the abstract page?
Not always. Many servers display contact data stored in page metadata, separate from the PDF. You often need both a new upload and a support request to edit the landing page.
Can I delete an old preprint version?
Policies vary. Some platforms rarely delete versions for record integrity. You can usually add a new version, request landing-page redactions, and ask to limit indexing of sensitive past versions.
Do I need to remove ORCID links?
No. ORCID is helpful for attribution. Instead, adjust visibility so your email and precise employment details aren’t public if you prefer privacy.
What about citations that already reference the old version?
They typically persist, but you can ensure the landing page emphasizes the updated version and includes a visible “latest version” link. Update your own profiles and lab pages to point to the clean version.
Will this affect peer review or future publication?
Redacting personal contact details from preprints generally doesn’t harm publication prospects. Journals prefer corresponding contact, but it can be a role-based address.
Conclusion
Preprint servers and author contribution notes help science move faster but can inadvertently expose sensitive personal details. The most effective approach is twofold: sanitize your files and metadata, then coordinate with the hosting platform to update the landing page and constrain legacy exposure. Follow a consistent checklist, keep your corresponding contact professional and non-personal, and monitor for reappearances across mirrors and indexes. With a few careful updates, you can keep your scholarship visible while keeping your personal information private.
Good to Know
Most preprint servers preserve version history even after you update files. You often need to both upload a new, redacted version and file a separate support request to hide or scrub the old version’s exposed details.