File Filter Tool by Text Character Count
Filter files that meet the criteria from multiple document formats (TXT, PDF, DOCX, PPTX, etc.) based on the set Chinese character count range, and support packaged downloads.
Browser execution mode: Your data is processed in your browser and is not uploaded to the server.
Speed and Stability: Processing speed depends on your device and browser. For large batch work, the desktop version may be more stable.
Loading tool, please wait...
Loading tool, please wait...
If the online tool fails to load or run, try the desktop tool.https://tools.yikeaigc.com/
Tool Usage
Return to old version
Confirm word range, counting method, and parsing options in “Filter Settings” before starting.
Files are read, counted, and packaged only on this page. For batch selection, the UI shows the first 20 files; all selected files are processed.
1Choose files or folder
Click to choose files, or drag documents here
Supports TXT, SRT, HTML, Markdown, DOCX, PDF, PPTX, XLSX, and more. You can also select a folder for batch reading.
When selecting a folder, supported documents are read recursively. To keep the UI tidy, only the first 20 are previewed.
2Sample Data
Selected Files
0
Total Size
0 KB
Supported Formats
0
Preview Count
0
3File Preview
No files selected. You can load sample data first, then go to filter settings.
| No. | File Name | Format | Size | Source Path |
|---|
1Word Count Range
Leave blank for no minimum.
Leave blank for no maximum.
The result ZIP includes only files matching the filter mode.
2Count Method
3Parsing & Cleanup
PDF only.
Enter 0 to process to the last page.
Results
After filtering, download a ZIP of matching files or export the full stats report as CSV.
After filtering, download a ZIP of matching files or export the full stats report as CSV.
Waiting0%
Process Files
0
Matched
0
No Match
0
Parse Failed
0
1Filter Results
Not processed yet. Go to “Filter Settings” to confirm parameters, then start processing.
| Select | File Name | Status | Stat Value | Chinese Characters | Format | Description |
|---|
2Process Log
Waiting to start.
Instructions
Software Usage Instructions
- Select files or folders: Click “Select Files” or “Select Folder”, or load the sample data directly. During batch import, the list only shows the first 20 files; the remaining files will be mentioned in the prompt, but all files will be included during processing.
- Set parameters first: Switch to “Filter Settings” and confirm the word count range, counting method, encoding method, PDF page range, and whitespace and punctuation handling rules in order before continuing to the next step.
- Start processing: After clicking “Start Processing”, the tool will read and count the file contents locally and will not upload them to the server.
- View results: After processing is complete, view the matched items, misses, and parsing failure lists on the results page. You can also switch the display by result range.
- Download output: Select the results you need and download them as ZIP or CSV. Result files will keep the original file names, and duplicate names will automatically have numbers appended for distinction.
- Continue batch processing: Each step page provides a “Next” button, making it easy to quickly jump to the next settings page and process again.
FAQ
A: Common formats such as TXT, SRT, MD, HTML, HTM, DOCX, PDF, PPTX, XLSX, CSV, JSON, XML, and XLS are supported, as well as recursive folder reading.
A: You can enter an exact range, such as “300-1200”; you can also enter a single number, which means a range from 0 to that value, suitable for filtering short text or empty content.
A: If you mainly focus on Chinese content, choose Chinese character count; if the document contains mixed Chinese and English, choose mixed Chinese-English; if you care more about the content actually displayed on the page, choose the visible character method.
A: When a PDF is very long, you can count only a specified page range to avoid including the entire document. This is useful when you only need to check a specific section.
A: A number is automatically appended to the original file name to distinguish it. For example, files with the same name become “filename(2)” and “filename(3)”, preventing duplicates within the ZIP archive.
A: Common reasons include file corruption, nonstandard formatting, encoding issues, empty content, or a mismatch between the file type and its extension. You can try again with a sample in a similar format first.
Related Tools
Batch Separate Compression Tool
Compress files or folders into ZIP files by project, with...
URL Batch Filtering and Processing Tool
Batch extract URLs from TXT text, filter and deduplicate ...
Batch File MD5 Modification Tool
Batch modify file MD5 values, generating different MD5 va...