Epigenetic specifications of host chromosome docking sites for latent Epstein-Barr virus

Epstein-Barr virus (EBV) genomes persist in latently infected cells as extrachromosomal episomes that attach to host chromosomes through the tethering functions of EBNA1, a viral encoded sequence-specific DNA binding protein. Here we employ circular chromosome conformation capture (4C) analysis to identify genome-wide associations between EBV episomes and host chromosomes. We find that EBV episomes in Burkitt’s lymphoma cells preferentially associate with cellular genomic sites containing EBNA1 binding sites enriched with B-cell factors EBF1 and RBP-jK, the repressive histone mark H3K9me3, and AT-rich flanking sequence. These attachment sites correspond to transcriptionally silenced genes with GO enrichment for neuronal function and protein kinase A pathways. Depletion of EBNA1 leads to a transcriptional de-repression of silenced genes and reduction in H3K9me3. EBV attachment sites in lymphoblastoid cells with different latency type show different correlations, suggesting that host chromosome attachment sites are functionally linked to latency type gene expression programs.


Statistics
For all statistical analyses, confirm that the following items are present in the figure legend, table legend, main text, or Methods section.

n/a Confirmed
The exact sample size (n) for each experimental group/condition, given as a discrete number and unit of measurement A statement on whether measurements were taken from distinct samples or whether the same sample was measured repeatedly The statistical test(s) used AND whether they are one-or two-sided Only common tests should be described solely by name; describe more complex techniques in the Methods section.
A description of all covariates tested A description of any assumptions or corrections, such as tests of normality and adjustment for multiple comparisons A full description of the statistical parameters including central tendency (e.g. means) or other basic estimates (e.g. regression coefficient) AND variation (e.g. standard deviation) or associated estimates of uncertainty (e.g. confidence intervals) For null hypothesis testing, the test statistic (e.g. F, t, r) with confidence intervals, effect sizes, degrees of freedom and P value noted Give P values as exact values whenever suitable.

For Bayesian analysis, information on the choice of priors and Markov chain Monte Carlo settings
For hierarchical and complex designs, identification of the appropriate level for tests and full reporting of outcomes Estimates of effect sizes (e.g. Cohen's d, Pearson's r), indicating how they were calculated Our web collection on statistics for biologists contains articles on many of the points above.

Software and code
Policy information about availability of computer code Data collection

Data analysis
For manuscripts utilizing custom algorithms or software that are central to the research but not yet described in published literature, software must be made available to editors/reviewers. We strongly encourage code deposition in a community repository (e.g. GitHub). See the Nature Research guidelines for submitting code & software for further information.

Data
Policy information about availability of data All manuscripts must include a data availability statement. This statement should provide the following information, where applicable: -Accession codes, unique identifiers, or web links for publicly available datasets -A list of figures that have associated raw data -A description of any restrictions on data availability Field-specific reporting Please select the one below that is the best fit for your research. If you are not sure, read the appropriate sections before making your selection.

Life sciences Behavioural & social sciences Ecological, evolutionary & environmental sciences
For a reference copy of the document with all sections, see nature.com/documents/nr-reporting-summary-flat.pdf Double-blind peer review submissions: write DBPR and your manuscript number here instead of author names.

YYYY-MM-DD
No software were used for data collection. Life sciences study design All studies must disclose on these points even when the disclosure is negative. anti-EBNA1 antibody is validated with shEBNA1 (Fig. 6a). Describe the authentication procedures for each cell line used OR declare that none of the cell lines used were authenticated.

MutuI (EBV
All cell lines tested negative for mycoplasma contamination.
Name any commonly misidentified cell lines used in the study and provide a rationale for their use.
Provide provenance information for specimens and describe permits that were obtained for the work (including the name of the issuing authority, the date of issue, and any identifying information).
Indicate where the specimens have been deposited to permit free access by other researchers.
If new dates are provided, describe how they were obtained (e.g. collection, storage, sample pretreatment and measurement), where they were obtained (i.e. lab name), the calibration program and the protocol for quality assurance OR state that no new dates are provided.

nature research | reporting summary
October 2018

Animals and other organisms
Policy information about studies involving animals; ARRIVE guidelines recommended for reporting animal research

Laboratory animals
Wild animals

Field-collected samples
Ethics oversight Note that full information on the approval of the study protocol must also be provided in the manuscript.

Human research participants
Policy information about studies involving human research participants Population characteristics

Recruitment
Ethics oversight Note that full information on the approval of the study protocol must also be provided in the manuscript.

Clinical data Policy information about clinical studies
All manuscripts should comply with the ICMJEguidelines for publication of clinical research and a completedCONSORT checklist must be included with all submissions.

Clinical trial registration
Study protocol

Data collection
Outcomes ChIP-seq Data deposition Confirm that both raw and final processed data have been deposited in a public database such as GEO.
Confirm that you have deposited or provided access to graph files (e.g. BED files) for the called peaks.

Data access links
May remain private before publication.

Files in database submission
Genome browser session (e.g. UCSC)

Methodology
Replicates Sequencing depth

Antibodies
For laboratory animals, report species, strain, sex and age OR state that the study did not involve laboratory animals.
Provide details on animals observed in or captured in the field; report species, sex and age where possible. Describe how animals were caught and transported and what happened to captive animals after the study (if killed, explain why and describe method; if released, say where and when) OR state that the study did not involve wild animals.
For laboratory work with field-collected samples, describe all relevant parameters such as housing, maintenance, temperature, photoperiod and end-of-experiment protocol OR state that the study did not involve samples collected from the field.
Identify the organization(s) that approved or provided guidance on the study protocol, OR state that no ethical approval or guidance was required and explain why not.
Describe the covariate-relevant population characteristics of the human research participants (e.g. age, gender, genotypic information, past and current diagnosis and treatment categories). If you filled out the behavioural & social sciences study design questions and have nothing to add here, write "See above." Describe how participants were recruited. Outline any potential self-selection bias or other biases that may be present and how these are likely to impact results.