Improper Validation of Integrity Check Valuenltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Improper Validation of Integrity Check Value in nltk/downloader.py through the Downloader._download_package() download-and-extract flow. An attacker can install malicious package contents by tampering with the archive response, such as through a compromised mirror, HTTP interception, or a substituted download source. The downloader writes the fetched package to disk and then proceeds to extract it without verifying the final file against the expected checksum, so a modified archive can be unpacked as if it were trusted. This can lead to attacker-controlled corpus or model files being installed on systems that use nltk.download().
How to fix Improper Validation of Integrity Check Value? Upgrade nltk to version 3.10.0 or higher.
| |
External Control of File Name or Pathnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to External Control of File Name or Path via the Downloader.download and Downloader.incr_download file-write path in nltk/downloader.py. An attacker who can place a hardlink inside a writable shared downloader root can make a normal package install overwrite the linked inode, mutating files outside the intended install tree.
How to fix External Control of File Name or Path? Upgrade nltk to version 3.10.3 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal due to raw file I/O in the TransitionParser, AveragedPerceptron, and PerceptronTagger model-artifact APIs. An attacker can read or overwrite files outside the allowed roots by supplying a model path that those APIs pass to ordinary open() calls even when pathsec enforcement is enabled. This breaks local containment for applications that rely on NLTK path security to keep model loads and saves inside a sandbox. It can expose sensitive files or clobber arbitrary files outside the approved directory tree through normal model import and export workflows.
Workarounds
- Use
pathsec enforcement only for trusted model paths, and do not let untrusted workflows choose model import or export destinations; otherwise TransitionParser, AveragedPerceptron, and PerceptronTagger can read or overwrite files outside the allowed roots.
How to fix Directory Traversal? Upgrade nltk to version 3.10.3 or higher.
| |
Uncontrolled Recursionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Uncontrolled Recursion via the RecursiveDescentParser in nltk/parse/recursivedescent.py. An attacker can pin a CPU core indefinitely or exhaust the Python recursion stack by supplying a tiny left-recursive or highly ambiguous grammar and triggering top-down parsing on even a short input. The parser enumerates parse trees recursively with no default bound on the search, so crafted grammars cause unbounded work and hang the process. This can deny service to applications that parse untrusted grammars or user-controlled input with RecursiveDescentParser or SteppingRecursiveDescentParser.
How to fix Uncontrolled Recursion? Upgrade nltk to version 3.10.3 or higher.
| |
Inefficient Algorithmic Complexitynltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Inefficient Algorithmic Complexity through the XMLCorpusView._read_xml_fragment() function in nltk/corpus/reader/xmldocs.py. An attacker can exhaust CPU by supplying a malformed XML corpus file with a single oversized unterminated tag to an affected reader such as BNCCorpusReader.words(). The function reads the file in 1 KiB blocks and repeatedly rescans the entire growing fragment with _VALID_XML_RE.match() and fragment.rfind("<"), so processing time grows quadratically with input size. A user opening attacker-controlled corpus data sees excessive CPU consumption and a denial of service while the reader tries to parse the file.
How to fix Inefficient Algorithmic Complexity? Upgrade nltk to version 3.10.3 or higher.
| |
Unchecked Input for Loop Conditionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Unchecked Input for Loop Condition via the FeatStructReader in nltk/featstruct.py. An attacker can trigger an unhandled RecursionError and crash parsing code by supplying a deeply nested feature-structure string to FeatStruct(s), FeatStructReader().fromstring(s), or FeatureGrammar.fromstring(s). The parser’s mutually recursive value-reading path has no nesting-depth limit, so each additional nested [...], {...}, or (...) pushes another Python stack frame until the interpreter’s recursion limit is exceeded. Applications that accept untrusted feature-structure or feature-grammar text can be taken down with a small crafted input, causing denial of service.
How to fix Unchecked Input for Loop Condition? Upgrade nltk to version 3.10.3 or higher.
| |
Inefficient Algorithmic Complexitynltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Inefficient Algorithmic Complexity via the PorterStemmer.stem() path in nltk/stem/porter.py. An attacker can pin a CPU core by supplying a single token with a long run of y characters, especially one that reaches stemming logic such as a y...ness input. The vulnerable code calls _is_consonant(word, i) once per character in _measure() and _contains_vowel(), and _is_consonant() walks backward across the same y run on every call. That makes stemming a long y token consume O(n^2) time, delaying or stalling requests that process untrusted text.
How to fix Inefficient Algorithmic Complexity? Upgrade nltk to version 3.10.3 or higher.
| |
Inefficient Algorithmic Complexitynltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Inefficient Algorithmic Complexity in the read_block method of TEICorpusView (nltk/corpus/reader/pl196x.py), whose lazy .*? whole-block regexes rescan the text from each opening-tag position. An attacker can force quadratic CPU growth and stall the parser thread by supplying a PL196X or TEI-like corpus file with many unmatched opening tags, so each regex attempt scans to the block end, fails, and restarts from the next tag. This requires the application to parse an attacker-influenced corpus file through the Pl196xCorpusReader public APIs such as words() or tagged_words().
How to fix Inefficient Algorithmic Complexity? Upgrade nltk to version 3.10.3 or higher.
| |
Arbitrary Argument Injectionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Arbitrary Argument Injection through the per-call options handling in the java() function in nltk/internals.py. An attacker can achieve arbitrary code execution by supplying crafted java_options containing malicious JVM flags to GenericStanfordParser, StanfordTagger, StanfordTokenizer, or StanfordSegmenter; the options are passed to the Java process without validation, allowing native or Java agents and other code-loading mechanisms to execute. This affects deployments where wrapper options are derived from untrusted user input, configuration, or environment data.
How to fix Arbitrary Argument Injection? Upgrade nltk to version 3.10.3 or higher.
| |
Regular Expression Denial of Service (ReDoS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Regular Expression Denial of Service (ReDoS) in TokenSearcher.findall() and Text.findall() in nltk/text.py. An attacker can pin a Python worker and exhaust CPU by supplying a crafted token-regex query, such as a quantified pattern that forces catastrophic backtracking when matched against a long token corpus. Text.findall() delegates directly to TokenSearcher.findall(), which applies the user-controlled expression across the joined corpus string with no execution bound. In applications that expose these search helpers to external input, a single malicious request can stall concordance or token search for other users until the process is interrupted.
How to fix Regular Expression Denial of Service (ReDoS)? Upgrade nltk to version 3.10.0 or higher.
| |
Regular Expression Denial of Service (ReDoS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Regular Expression Denial of Service (ReDoS) through the _tgrep_node_action() path in nltk/tgrep.py. An attacker can pin a CPU core indefinitely by supplying a tgrep query with a crafted /regex/ node literal that is compiled and searched against tree node labels without a time bound. This affects applications that expose tgrep_positions() or tgrep_compile() to external input, where a single malicious pattern can hang the Python process and block service for other users.
How to fix Regular Expression Denial of Service (ReDoS)? Upgrade nltk to version 3.10.3 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal through IPIPANCorpusReader, CrubadanCorpusReader, and LinThesaurusCorpusReader in nltk/corpus/reader/ipipan.py, nltk/corpus/reader/crubadan.py, and nltk/corpus/reader/lin.py. An attacker can disclose files outside a trusted corpus root by placing symlinks inside that root and triggering the readers’ normal corpus-loading methods.
How to fix Directory Traversal? Upgrade nltk to version 3.10.3 or higher.
| |
Deserialization of Untrusted Datanltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Deserialization of Untrusted Data in nltk.picklesec.AllowlistUnpickler, nltk.tokenize.punkt.punkt_pickle_load, and nltk.parse.transitionparser.TransitionParser.parse. An attacker can execute arbitrary commands by supplying a crafted pickle that uses REDUCE to reach dangerous callables inside an allowlisted namespace, such as numpy.f2py.crackfortran.myeval or nltk.tokenize.repp.ReppTokenizer._execute. This lets a malicious tokenizer or model artifact run code during unpickling, so applications that load untrusted NLTK pickles can be compromised while believing the loader is safe.
How to fix Deserialization of Untrusted Data? Upgrade nltk to version 3.10.3 or higher.
| |
External Control of File Name or Pathnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to External Control of File Name or Path through the CorpusReader.__init__ constructor and the LinThesaurusCorpusReader and PanLexLiteCorpusReader implementations in nltk/corpus/reader/api.py, nltk/corpus/reader/lin.py, and nltk/corpus/reader/panlex_lite.py. An attacker can make NLTK read files or open an SQLite database outside the intended nltk.pathsec sandbox by supplying a corpus root path to a reader constructor. The vulnerable constructors turn caller-supplied roots into paths and then reach raw open() or sqlite3.connect() on derived locations without sandbox validation, so a process using these readers can expose out-of-root corpus data. In the affected deployments, this lets untrusted corpus-root input bypass the data-root boundary and load local file contents from outside the intended corpus tree.
How to fix External Control of File Name or Path? Upgrade nltk to version 3.10.3 or higher.
| |
Deserialization of Untrusted Datanltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Deserialization of Untrusted Data through the parse() method in nltk/parse/transitionparser.py. An attacker can execute arbitrary Python code by supplying a malicious parser model file that TransitionParser.parse() loads with pickle_load(f). This affects applications that accept or load untrusted transition parser model paths, letting the payload run with the privileges of the process loading the model.
How to fix Deserialization of Untrusted Data? Upgrade nltk to version 3.10.0 or higher.
| |
Untrusted Search Pathnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Untrusted Search Path through the dot2img function in nltk/parse/dependencygraph.py and the AlignedSent._repr_svg_ method in nltk/translate/api.py. An attacker can execute an arbitrary planted dot binary by supplying a malicious dot on the process search path or in the current working directory when these Graphviz rendering paths invoke dot by bare name. This can lead to arbitrary code execution in the context of the process that renders dependency graphs or SVG output, causing user systems to run attacker-controlled code instead of the intended Graphviz binary.
How to fix Untrusted Search Path? Upgrade nltk to version 3.10.3 or higher.
| |
Improper Restriction of Recursive Entity References in DTDs ('XML Entity Expansion')nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Improper Restriction of Recursive Entity References in DTDs ('XML Entity Expansion') via the raw xml.etree.ElementTree parsing in nltk.chunk.named_entity.load_ace_file, nltk.internals.ElementWrapper, and nltk.downloader’s XML loaders. An attacker can exhaust memory by supplying a crafted XML document with nested <!ENTITY> declarations that expand a few hundred bytes into megabytes when NLTK parses ACE annotation files, package metadata, or arbitrary XML strings.
How to fix Improper Restriction of Recursive Entity References in DTDs ('XML Entity Expansion')? Upgrade nltk to version 3.10.3 or higher.
| |
Server-side Request Forgery (SSRF)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Server-side Request Forgery (SSRF) via the urlopen process in nltk/pathsec.py. An attacker can reach internal-only HTTP resources by supplying a URL that passes local validation and then fetching it through a configured proxy, which performs the actual egress outside NLTK’s pinning path. This lets an attacker read loopback, link-local, or private network content through callers such as nltk.data.load and downloader fetch paths, causing internal data exposure and allowing the application to retrieve and process forged downloader content.
How to fix Server-side Request Forgery (SSRF)? Upgrade nltk to version 3.10.3 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal in StreamBackedCorpusView that bypasses pathsec.ENFORCE by calling builtins.open() directly instead of pathsec.open(). Attackers who control the fileid argument can read arbitrary local files regardless of the ENFORCE setting, including sensitive system files and application credentials.
How to fix Directory Traversal? Upgrade nltk to version 3.10.0 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal via the FileSystemPathPointer.open function. An attacker can access arbitrary files accessible to the process user by supplying crafted file:// URLs to the nltk.data.load function.
How to fix Directory Traversal? Upgrade nltk to version 3.10.0 or higher.
| |
External Control of File Name or Pathnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to External Control of File Name or Path via path handling in FramenetCorpusReader and NKJPCorpusReader. An attacker can read and parse arbitrary XML files outside the configured corpus root by supplying crafted selectors or poisoning index state. Attacker-controlled values passed through methods such as frame_by_name(), doc(), lu(), and header() can traverse outside the intended directory and access XML files readable by the application.
How to fix External Control of File Name or Path? Upgrade nltk to version 3.10.0 or higher.
| |
Insecure Default Initialization of Resourcenltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Insecure Default Initialization of Resource via disabled security enforcement in pathsec.py. An attacker can bypass protections against path traversal and unsafe pickle deserialization because NLTK defaults ENFORCE to False. In this configuration, failed security validation checks only generate warnings instead of rejecting malicious input, allowing operations that violate the intended security restrictions to continue.
How to fix Insecure Default Initialization of Resource? Upgrade nltk to version 3.10.0 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal via the _get_tag() method in nltk/corpus/reader/ipipan.py, which is the single choke point for channels(), domains(), categories(), and fileids(channels=...). These methods call .replace("morph.xml", "header.xml") on the result of self.abspath(), and because FileSystemPathPointer subclasses str, the .replace() call silently returns a plain string, discarding the PathPointer wrapper. The resulting open() call therefore bypasses nltk.pathsec entirely — including the symlink-resolving, root-scoped containment check used elsewhere in the library. An attacker who can plant a symlink inside the corpus root (with a name containing no path separators or .., so it passes the existing traversal guard) can point it to any file readable by the process, allowing arbitrary local file disclosure.
How to fix Directory Traversal? Upgrade nltk to version 3.10.2 or higher.
| |
Symlink Attacknltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Symlink Attack via symlink handling in CorpusReader.open(). A local attacker can read arbitrary files outside the configured corpus root by placing symbolic links inside the corpus directory that resolve to files outside the intended boundary. Because path validation operates on the lexical path without validating the resolved symlink target, the symlink passes the directory restriction and is subsequently followed during file access.
How to fix Symlink Attack? Upgrade nltk to version 3.9.4 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal via symlink handling in FramenetCorpusReader. An attacker can read arbitrary XML files outside the configured corpus root by placing symlinks with separator-free names inside the corpus directory. These names pass the path validation checks, but resolve to targets outside the intended sandbox when accessed through methods such as frame_by_name(), _lu_file(), or doc().
How to fix Directory Traversal? Upgrade nltk to version 3.10.2 or higher.
| |
Regular Expression Denial of Service (ReDoS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Regular Expression Denial of Service (ReDoS) via regular expression processing in TweetTokenizer.tokenize() and casual_tokenize(). An attacker can cause denial of service by supplying crafted text containing repeated domain-label separators that triggers catastrophic backtracking in the URLS regular expression. Because the affected domain-matching branch can be partitioned in exponentially many ways before ultimately failing, relatively small inputs can consume seconds or minutes of CPU and stall services that tokenize untrusted text.
How to fix Regular Expression Denial of Service (ReDoS)? Upgrade nltk to version 3.10.3 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal via multiple corpus readers (crubadan.py, lin.py, xmldocs.py, pl196x.py, mte.py, toolbox.py, named_entity.py, nkjp.py, ipipan.py) that opened files using raw open() or codecs.open() calls, bypassing NLTK's trusted-root containment. An attacker who can place a symlink or hardlink inside a trusted corpus root directory can cause the affected readers to follow the link and read files outside the intended root, disclosing their contents through normal corpus reader methods. Exploitation requires the ability to write a symlink or hardlink into a corpus directory that is subsequently read by the application.
How to fix Directory Traversal? Upgrade nltk to version 3.10.3 or higher.
| |
Server-side Request Forgery (SSRF)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Server-side Request Forgery (SSRF) via the validate_network_url() DNS resolution path in nltk/pathsec.py. An attacker can make urlopen() reach restricted internal URLs by causing hostname resolution to fail, which leaves the validation loop with no resolved addresses to check and lets the request proceed without rejecting the target. When DNS is unavailable or returns an error for an attacker-supplied hostname, the SSRF filter fails open instead of blocking the request. This can expose internal services such as cloud metadata endpoints to network requests initiated by NLTK and bypass URL validation in applications that rely on it.
How to fix Server-side Request Forgery (SSRF)? Upgrade nltk to version 3.10.0 or higher.
| |
Server-side Request Forgery (SSRF)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Server-side Request Forgery (SSRF) via the validate_network_url() address-policy check in nltk/pathsec.py. An attacker can make a strict-mode application send requests to shared-address-space or other non-global hosts by supplying a URL whose resolved IP falls in 100.64.0.0/10, which the filter does not reject. The vulnerable logic only blocked loopback, link-local, multicast, and private addresses, so requests to RFC 6598 carrier-grade NAT targets passed validation. That exposes internal or non-public network resources reachable from the host running NLTK’s network-loading helpers.
How to fix Server-side Request Forgery (SSRF)? Upgrade nltk to version 3.10.0 or higher.
| |
Time-of-check Time-of-use (TOCTOU) Race Conditionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Time-of-check Time-of-use (TOCTOU) Race Condition via the downloader process. An attacker can compromise data integrity by causing package archives to be extracted into shared namespaces before integrity validation, which may overwrite trusted resources in machine learning pipelines and environments sensitive to reproducibility. This is only exploitable if a user interacts with the downloader and downloads a malicious package.
How to fix Time-of-check Time-of-use (TOCTOU) Race Condition? There is no fixed version for nltk.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal via FramenetCorpusReader.frame() in nltk/corpus/reader/framenet.py. An attacker can read arbitrary XML files by supplying a frame name containing ../ sequences, causing the reader to escape the corpus root and return the parsed file contents. The vulnerable code interpolates the caller-supplied frame name directly into the frame/<name>.xml path and passes that string to XMLCorpusView, bypassing the CorpusReader.open() and nltk.pathsec sandbox checks, including ENFORCE=True. This can disclose out-of-corpus XML data to any application that exposes FrameNet frame lookups to untrusted input.
How to fix Directory Traversal? Upgrade nltk to version 3.10.0 or higher.
| |
Regular Expression Denial of Service (ReDoS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Regular Expression Denial of Service (ReDoS) via the FEATURES regex in nltk/corpus/reader/reviews.py. An attacker can hang ReviewsCorpusReader.reviews(), features(), or sents() by supplying a reviews corpus line made of a long run of words with no trailing bracketed feature annotation. The unbounded feature-label pattern backtracks quadratically on such input, burning CPU in the reader and stalling applications that process untrusted or user-supplied review corpora.
How to fix Regular Expression Denial of Service (ReDoS)? Upgrade nltk to version 3.10.0 or higher.
| |
Server-side Request Forgery (SSRF)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Server-side Request Forgery (SSRF) through the urlopen process in nltk.pathsec. An attacker can reach loopback or internal HTTP services by supplying a hostname that resolves to a public IP during SSRF validation and to a private IP when urllib makes the actual connection. This defeats nltk.pathsec.ENFORCE and lets nltk.download(server_index_url=...) or nltk.data.load("http://...") fetch and return content from services that the filter was meant to block, including internal admin endpoints and metadata services.
How to fix Server-side Request Forgery (SSRF)? Upgrade nltk to version 3.10.0 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal via the add_root path handling in nltk/corpus/reader/nkjp.py. An attacker can read files outside the corpus root and bypass the nltk.pathsec sandbox by supplying fileids that contain .. sequences, absolute paths, or symlink escapes to the header, raw, words, sents, or tagged_words read methods. The reader concatenates the caller-controlled fileids into filesystem paths and opens them with the builtin open(), so a crafted path can reach an arbitrary header.xml or other target under attacker-chosen directories on the host. In deployments that rely on ENFORCE=True, this exposes local files that were supposed to be blocked by NLTK’s path validation.
How to fix Directory Traversal? Upgrade nltk to version 3.10.0 or higher.
| |
Eval Injectionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Eval Injection via the __main__ process in collocations.py when command-line arguments are passed directly to the eval function without proper validation or sanitization. An attacker can execute arbitrary Python code, including operating system commands, by supplying crafted command-line arguments.
How to fix Eval Injection? Upgrade nltk to version 3.9.3 or higher.
| |
Improper Verification of Cryptographic Signaturenltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Improper Verification of Cryptographic Signature via the Stanford Interface Classes. An attacker can execute arbitrary code by supplying a malicious JAR file to the affected Stanford interface classes, which are invoked without integrity verification.
How to fix Improper Verification of Cryptographic Signature? Upgrade nltk to version 3.10.3 or higher.
| |
Missing Authentication for Critical Functionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Missing Authentication for Critical Function via the nltk.app.wordnet_app module. An attacker can cause the server to shut down and disrupt service availability by sending a specially crafted GET request to the WordNet Browser HTTP server when it is running in its default mode.
How to fix Missing Authentication for Critical Function? Upgrade nltk to version 3.9.4 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal via the nltk.data.load() function. An attacker can access arbitrary files on the local filesystem by supplying specially crafted URL-encoded path traversal payloads that bypass input validation and are decoded after security checks.
How to fix Directory Traversal? Upgrade nltk to version 3.10.0 or higher.
| |
Unsafe Dependency Resolutionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Unsafe Dependency Resolution due to lack of verification or sandboxing in the StanfordSegmenter module, when unvalidated Java Archive (JAR) files are dynamically loaded. An attacker can execute arbitrary Java bytecode by supplying or replacing a JAR file, potentially through model poisoning, Man-in-the-Middle (MITM) attacks, or dependency poisoning.
How to fix Unsafe Dependency Resolution? Upgrade nltk to version 3.9.3 or higher.
| |
Uncontrolled Recursionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Uncontrolled Recursion via the JSONTaggedDecoder.decode_obj() function in jsontags.py. An attacker can cause the application to crash by submitting a deeply nested JSON structure that exceeds the recursion limit, resulting in an unhandled exception.
How to fix Uncontrolled Recursion? Upgrade nltk to version 3.9.4 or higher.
| |
Cross-site Scripting (XSS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Cross-site Scripting (XSS) via the lookup_... route in the web interface, where attacker-controlled input is reflected into the HTML response without proper escaping. An attacker can execute arbitrary JavaScript in the browser context of the application by convincing a user to open a crafted URL.
How to fix Cross-site Scripting (XSS)? Upgrade nltk to version 3.9.4 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal via the XML index file downloader. An attacker can overwrite arbitrary files and create directories at unintended locations by supplying malicious values for the subdir and id attributes containing path traversal sequences in a remote XML index file.
How to fix Directory Traversal? Upgrade nltk to version 3.9.4 or higher.
| |
Missing Authentication for Critical Functionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Missing Authentication for Critical Function in WordNet Browser HTTP server in default configuration. An attacker can cause the service to terminate immediately by sending a specially crafted unauthenticated HTTP GET request (e.g. http://127.0.0.1:8004/SHUTDOWN%20THE%20SERVER) to the listening port.
How to fix Missing Authentication for Critical Function? Upgrade nltk to version 3.9.4 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal via the filestring function. An attacker can access sensitive files by supplying specially crafted input paths, such as absolute paths or directory traversal sequences, to bypass input validation and read arbitrary files on the system.
Note:
This is only exploitable if the function processes untrusted user input, such as in web APIs or network-accessible services.
How to fix Directory Traversal? Upgrade nltk to version 3.9.3 or higher.
| |
Directory Traversalnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Directory Traversal due to improper sanitization of file paths in the CorpusReader classes. An attacker can gain unauthorized access to sensitive files by supplying crafted file paths to applications that process user-controlled file inputs.
How to fix Directory Traversal? Upgrade nltk to version 3.9.3 or higher.
| |
Arbitrary Code Injectionnltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Arbitrary Code Injection via the _unzip_iter() function due to the lack of validation before unpacking untrusted downloaded packages. An attacker can execute arbitrary code by supplying a specially crafted zip file.
How to fix Arbitrary Code Injection? Upgrade nltk to version 3.9.3 or higher.
| |
Remote Code Execution (RCE)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Remote Code Execution (RCE) through the integrated data package download functionality. An attacker with control over the NLTK data index can execute arbitrary code by supplying pickled Python code within untrusted packages and trick a user into loading the malicious pickle.
Some packages found to be vulnerable if compromised are averaged_perceptron_tagger, punkt, maxent_ne_chunker, help/tagsets, and maxent_treebank_pos_tagger.
How to fix Remote Code Execution (RCE)? Upgrade nltk to version 3.9 or higher.
| |
Remote Code Execution (RCE)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Remote Code Execution (RCE) in the local WordNet browser. When a user opens a malicious link while the WordNet browser is active, it can result in the exploitation of this vulnerability on their system.
How to fix Remote Code Execution (RCE)? Upgrade nltk to version 3.8.1 or higher.
| |
Cross-site Scripting (XSS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Cross-site Scripting (XSS) due to improper input sanitization in the local Wordnet browser via the MyServerHandler class. Exploiting this vulnerability is possible by creating a maliciously crafted URL.
Note:
This only affects users of this browser interface to Wordnet, and not other users of Wordnet.
How to fix Cross-site Scripting (XSS)? Upgrade nltk to version 3.8.1 or higher.
| |
Regular Expression Denial of Service (ReDoS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Regular Expression Denial of Service (ReDoS) via the RegexpTagger method.
How to fix Regular Expression Denial of Service (ReDoS)? Upgrade nltk to version 3.6.6 or higher.
| |
Regular Expression Denial of Service (ReDoS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Regular Expression Denial of Service (ReDoS) via word_tokenize() in nltk/tokenize/punkt.py.
PoC
from nltk.tokenize import word_tokenize
import nltk
nltk.download('punkt')
import time
for length in [1000*2**n for n in range(1000)]:
text = "a" * length
start_t = time.time()
word_tokenize(text)
print(f"payload length: {length} takes {time.time()-start_t}s")
How to fix Regular Expression Denial of Service (ReDoS)? Upgrade nltk to version 3.6.6 or higher.
| |
Regular Expression Denial of Service (ReDoS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Regular Expression Denial of Service (ReDoS) in the CorpusReader for the Comparative Sentences Dataset.
How to fix Regular Expression Denial of Service (ReDoS)? Upgrade nltk to version 3.6.4 or higher.
| |
Regular Expression Denial of Service (ReDoS)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Regular Expression Denial of Service (ReDoS). _XML_TAG_NAME regex operator is vulnerable mainly due to the sub-pattern \s*/?\s* and can be exploited with an input such as "<"+" " * 5000
How to fix Regular Expression Denial of Service (ReDoS)? Upgrade nltk to version 3.6 or higher.
| |
Arbitrary File Write via Archive Extraction (Zip Slip)nltk is a Natural Language Toolkit (NLTK) is a Python package for natural language processing.
Affected versions of this package are vulnerable to Arbitrary File Write via Archive Extraction (Zip Slip).
It allows attackers to write arbitrary files via a ../ (dot dot slash) in an NLTK package (ZIP archive) that is mishandled during extraction.
How to fix Arbitrary File Write via Archive Extraction (Zip Slip)? Upgrade nltk to version 3.4.5 or higher.
| |