China Open-Government Data Download
Budget / Salary€30–250
TypeFreelance project
LocationRemote
Posted1 hour ago
Title: Download public open-government datasets from Chinese provincial open-data portals (China-based, real-name account required)
Budget: $250 USD, fixed-price, milestone-based (pilot first). Project 1 of a planned series — reliable delivery leads to further, larger projects.
Summary
We need current snapshots of specific publicly-published market-entity datasets from Chinese provincial/municipal government open-data platforms (公共数据开放平台). The datasets are openly published, but downloading the full file requires a registered real-name account (身份证 + 大陆手机号). We need a China-based freelancer with a verified real-name account to log in, download the target datasets in full, and deliver them as clean files + a metadata sheet. Data-retrieval through each portal's official open-data channel — not scraping, no CAPTCHA-solving, no non-public endpoints.
What we need — the data
For each portal, the market-entity / enterprise-registration master dataset (市场主体 / 企业登记 / 营业执照 / 商事主体 信息). A dataset qualifies only if its columns include all three:
- 统一社会信用代码 (18-char USCC)
- 企业名称 / 名称
- 登记状态 / 经营状态 / 企业状态 (e.g. 存续 / 注销 / 吊销)
Useful if present (take them): 成立日期, 法定代表人, 注册资本, 注册地址, 主体类型 (企业/个体工商户), 行业.
We want the full dataset matching the portal's stated row count (数据量) — not the anonymous preview/sample (anonymous download often caps at ~100,000 rows; the logged-in full file is the deliverable).
Freshness is critical — report the real last-update date
Outdated data is useless to us. For every dataset you must report, and we verify, the dataset's own last-update date — the portal's 最近更新时间 / 数据更新时间 field on the dataset page — which is different from both the stated 更新周期 (promised cycle) and your download date.
- Report all three: (a) 更新周期 cycle, (b) 最近更新时间 actual last data update, (c) download date.
- If the actual last-update date is inconsistent with the cycle (a "daily"/"monthly" set last really updated months/years ago), it is stale — flag and report it. Do not deliver stale data as if current.
- Include a screenshot of the dataset page showing 最近更新时间.
Access method
- Log in with your own real-name verified account.
- 无条件开放 (unconditional): download the full file directly after login.
- 有条件开放 (conditional): if a set is conditional, note it and tell us the application requirement — but for this project, prioritise the unconditional sets; do not wait on approvals.
- Prefer CSV / XLSX / JSON; note if only an API/接口 is available.
Portals in scope (this project)
- PILOT — Zhejiang 浙江 — data.zjzwfw.gov.cn — dataset: 市场监管市场主体信息(含个体) (has an unconditional copy, updated daily)
- Hunan 湖南 — data.hunan.gov.cn — dataset: 企业公示信息 / 市场主体
- Guangdong 广东 — gddata.gd.gov.cn — dataset: 市场主体 (per-city sets)
- Jiangsu 江苏 — data.jiangsu.gov.cn — dataset: 市场主体 (if the province portal has none, say so)
- Shenzhen 深圳 — opendata.sz.gov.cn — dataset: 个体工商户 + 商事主体基本信息
- Hangzhou 杭州 — data.hangzhou.gov.cn — dataset: 在册市场主体
- Wenzhou 温州 — data.wenzhou.gov.cn — dataset: 省市回流_市场主体信息
- Suzhou 苏州 — suzhou.gov.cn/dataOpenWeb — dataset: 营业执照照面信息 / 在业核验
Milestones
- Milestone 1 — Pilot (Zhejiang only): $40. Full Zhejiang dataset + metadata + screenshot. Released once it passes acceptance.
- Milestone 2 — remaining 7 portals: $210. Same standard.
Deliverables (per portal)
1. The dataset file(s), complete.
2. A metadata sheet (CSV/XLSX), one row per dataset: portal + URL, dataset name, dataset ID/链接, 开放类型, 更新周期, 最近更新时间, download date, rows delivered vs. rows stated, full column list, access method.
3. Screenshot(s) of each dataset page showing name, open-type, 更新周期, 最近更新时间, row count.
Acceptance criteria
- File parses; encoding stated (UTF-8 preferred).
- Has the three required columns (USCC + name + status).
- Row count within a small margin of the portal's stated 数据量 (the full set).
- 最近更新时间 reported and consistent with the cycle (stale sets flagged, not passed off as current).
- Metadata sheet complete and accurate.
Freelancer requirements
- Mainland China, real-name verified account for provincial 公共数据开放平台 / 政务服务网.
- Experience with government open-data portals a strong plus.
- Comfortable with large CSV/Excel files.
Scope boundaries
- Only officially published open datasets via each portal's official download/API channel, using a registered account.
- No scraping of non-open endpoints, no CAPTCHA circumvention, nothing paywalled/restricted.
- If a set isn't downloadable or is stale, report it — accuracy over workarounds.
To apply, tell us
1. Confirm you have a mainland real-name account and have downloaded open datasets from these portals before (which?).
2. For the Zhejiang pilot dataset: is it downloadable in full with your account right now, in what format(s), and what is its 最近更新时间 today? (Screenshot = priority.)
3. Any portal above where you expect a problem (no qualifying set / conditional-only / stale).
Budget: $250 USD, fixed-price, milestone-based (pilot first). Project 1 of a planned series — reliable delivery leads to further, larger projects.
Summary
We need current snapshots of specific publicly-published market-entity datasets from Chinese provincial/municipal government open-data platforms (公共数据开放平台). The datasets are openly published, but downloading the full file requires a registered real-name account (身份证 + 大陆手机号). We need a China-based freelancer with a verified real-name account to log in, download the target datasets in full, and deliver them as clean files + a metadata sheet. Data-retrieval through each portal's official open-data channel — not scraping, no CAPTCHA-solving, no non-public endpoints.
What we need — the data
For each portal, the market-entity / enterprise-registration master dataset (市场主体 / 企业登记 / 营业执照 / 商事主体 信息). A dataset qualifies only if its columns include all three:
- 统一社会信用代码 (18-char USCC)
- 企业名称 / 名称
- 登记状态 / 经营状态 / 企业状态 (e.g. 存续 / 注销 / 吊销)
Useful if present (take them): 成立日期, 法定代表人, 注册资本, 注册地址, 主体类型 (企业/个体工商户), 行业.
We want the full dataset matching the portal's stated row count (数据量) — not the anonymous preview/sample (anonymous download often caps at ~100,000 rows; the logged-in full file is the deliverable).
Freshness is critical — report the real last-update date
Outdated data is useless to us. For every dataset you must report, and we verify, the dataset's own last-update date — the portal's 最近更新时间 / 数据更新时间 field on the dataset page — which is different from both the stated 更新周期 (promised cycle) and your download date.
- Report all three: (a) 更新周期 cycle, (b) 最近更新时间 actual last data update, (c) download date.
- If the actual last-update date is inconsistent with the cycle (a "daily"/"monthly" set last really updated months/years ago), it is stale — flag and report it. Do not deliver stale data as if current.
- Include a screenshot of the dataset page showing 最近更新时间.
Access method
- Log in with your own real-name verified account.
- 无条件开放 (unconditional): download the full file directly after login.
- 有条件开放 (conditional): if a set is conditional, note it and tell us the application requirement — but for this project, prioritise the unconditional sets; do not wait on approvals.
- Prefer CSV / XLSX / JSON; note if only an API/接口 is available.
Portals in scope (this project)
- PILOT — Zhejiang 浙江 — data.zjzwfw.gov.cn — dataset: 市场监管市场主体信息(含个体) (has an unconditional copy, updated daily)
- Hunan 湖南 — data.hunan.gov.cn — dataset: 企业公示信息 / 市场主体
- Guangdong 广东 — gddata.gd.gov.cn — dataset: 市场主体 (per-city sets)
- Jiangsu 江苏 — data.jiangsu.gov.cn — dataset: 市场主体 (if the province portal has none, say so)
- Shenzhen 深圳 — opendata.sz.gov.cn — dataset: 个体工商户 + 商事主体基本信息
- Hangzhou 杭州 — data.hangzhou.gov.cn — dataset: 在册市场主体
- Wenzhou 温州 — data.wenzhou.gov.cn — dataset: 省市回流_市场主体信息
- Suzhou 苏州 — suzhou.gov.cn/dataOpenWeb — dataset: 营业执照照面信息 / 在业核验
Milestones
- Milestone 1 — Pilot (Zhejiang only): $40. Full Zhejiang dataset + metadata + screenshot. Released once it passes acceptance.
- Milestone 2 — remaining 7 portals: $210. Same standard.
Deliverables (per portal)
1. The dataset file(s), complete.
2. A metadata sheet (CSV/XLSX), one row per dataset: portal + URL, dataset name, dataset ID/链接, 开放类型, 更新周期, 最近更新时间, download date, rows delivered vs. rows stated, full column list, access method.
3. Screenshot(s) of each dataset page showing name, open-type, 更新周期, 最近更新时间, row count.
Acceptance criteria
- File parses; encoding stated (UTF-8 preferred).
- Has the three required columns (USCC + name + status).
- Row count within a small margin of the portal's stated 数据量 (the full set).
- 最近更新时间 reported and consistent with the cycle (stale sets flagged, not passed off as current).
- Metadata sheet complete and accurate.
Freelancer requirements
- Mainland China, real-name verified account for provincial 公共数据开放平台 / 政务服务网.
- Experience with government open-data portals a strong plus.
- Comfortable with large CSV/Excel files.
Scope boundaries
- Only officially published open datasets via each portal's official download/API channel, using a registered account.
- No scraping of non-open endpoints, no CAPTCHA circumvention, nothing paywalled/restricted.
- If a set isn't downloadable or is stale, report it — accuracy over workarounds.
To apply, tell us
1. Confirm you have a mainland real-name account and have downloaded open datasets from these portals before (which?).
2. For the Zhejiang pilot dataset: is it downloadable in full with your account right now, in what format(s), and what is its 最近更新时间 today? (Screenshot = priority.)
3. Any portal above where you expect a problem (no qualifying set / conditional-only / stale).
Apply on Freelancer →
Project sourced from Freelancer.com. Applications happen directly on the original platform — we never collect your data.